12 sep
|
Agileengine
|
Xico
DevOps / Site Reliability Engineer ID*****Full time | AgileEngine | MexicoPosted On 08/29/2026Job InformationCity Ciudad de MéxicoState/Province México*****IT ServicesJob DescriptionAgileEngine is an Inc. **** company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries.
We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.WHY JOIN USIf you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!
ABOUT THE ROLEWe are looking for a DevOps / Site Reliability Engineer to maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment.
This role blends platform engineering with incident command, using Terraform, CI/CD pipelines, and CSPM tools like Wiz.
You will lead major-incident calls, own remediation follow-through, and build the playbooks that guide response.WHAT YOU WILL DO- - Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP).
- - Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations.
- - Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment.
- - Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads.
- - Serve as Incident Commander on major and critical incidents — running the bridge, directing technical workstreams, making time‐critical decisions,
and coordinating cross‐functional responders under pressure.
- - Own the post‐incident loop — track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups.
- - Draft and send clear, accurate, audience‐appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle.
- - Develop, maintain, and socialize divisional / group‐level incident‐management playbooks, runbooks, and escalation procedures that standardize response and reduce time‐to‐resolution.MUST HAVES- - 5+ years of experience.
- - In-depth architectural expertise in multi‐cloud defense, federated IAM, and zero‐trust principles.
- - Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting.
- - Senior‐level, hands‐on incident‐command experience driving major/critical incident calls to resolution in a 24x7 production environment.
- - Proven track record of remediation follow‐up — coordinating with teams and holding owners accountable until issues are fully closed.
- - Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences.
- - Direct experience authoring divisional/group incident‐management playbooks and escalation procedures.
- - Fully autonomous.
- - Drives the architecture of complex automated runbooks and mentors Middle‐level SREs.
- - Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz.
- - Prior experience building platforms subject to strict financial compliance standards ( PCI‐DSS, SOC2 ).
NICE TO HAVES- - PagerDuty — hands‐on experience with on‐call scheduling, alert routing, and incident orchestration.
- - ServiceNow — familiarity with incident, problem, and change management workflows and reporting.PERKS AND BENEFITS- - Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget- - Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews- - Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm- - Meaningful, modern projects: build impactful products using modern technologies alongside integral teams and leading brands- - Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized- - Well‐being & support: access local well‐being programs and people‐focused support tailored to your location#J-*****-Ljbffr
📌 Devops / Site Reliability Engineer Id70127 (Xico)
🏢 Agileengine
📍 Xico