05 sep
|
Agileengine
|
Ciudad de México
05 sep
Agileengine
Ciudad de México
Dev Ops / Site Reliability Engineer ID70127 Full time | Agile Engine | Mexico Posted On 08/29/2026 Job Information City Ciudad de México State/Province México 01210 IT Services Job Description Agile Engine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards. WHY JOIN US If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you! ABOUT THE ROLE We are looking for a Dev Ops / Site Reliability Engineer to maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. This role blends platform engineering with incident command, using Terraform, CI/CD pipelines, and CSPM tools like Wiz. You will lead major-incident calls, own remediation follow-through, and build the playbooks that guide response. WHAT YOU WILL DO - Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP). - Engineer unified security policies and configuration baselines using Ia C (Terraform) to prevent misconfigurations. - Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment. - Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads. - Serve as Incident Commander on major and critical incidents — running the bridge, directing technical workstreams,
making time‑critical decisions, and coordinating cross‑functional responders under pressure. - Own the post‑incident loop — track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups. - Draft and send clear, accurate, audience‑appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle. - Develop, maintain, and socialize divisional / group‑level incident‑management playbooks, runbooks, and escalation procedures that standardize response and reduce time‑to‑resolution. MUST HAVES - 5+ years of experience . - In-depth architectural expertise in multi‑cloud defense, federated IAM, and zero‑trust principles . - Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting . - Senior‑level, hands‑on incident‑command experience driving major/critical incident calls to resolution in a 24x7 production environment . - Proven track record of remediation follow‑up — coordinating with teams and holding owners accountable until issues are fully closed.
- Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences. - Direct experience authoring divisional/group incident‑management playbooks and escalation procedures. - Fully autonomous. - Drives the architecture of complex automated runbooks and mentors Middle‑level SREs. - Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz . - Prior experience building platforms subject to strict financial compliance standards ( PCI‑DSS, SOC2 ). NICE TO HAVES - Pager Duty — hands‑on experience with on‑call scheduling, alert routing, and incident orchestration. - Service Now — familiarity with incident, problem, and change management workflows and reporting. PERKS AND BENEFITS - Growth without limits : build your skills through mentorship, internal Tech Talks, challenging projects, and a dedicated annual learning budget - Competitive compensation : get recognition that reflects your skills and impact, with regular performance and compensation reviews - Flexibility : work 100% remotely with adaptable hours that support focus, autonomy, and a healthy work rhythm - Meaningful, modern projects : build impactful products using modern technologies alongside global teams and leading brands - Collaborative culture : join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized - Well‑being & support : access local well‑being programs and people‑focused support tailored to your location #J-18808-Ljbffr
📌 Devops / site reliability engineer id70127 (Ciudad de México)
🏢 Agileengine
📍 Ciudad de México