29 ago
|
Agileengine
|
México
29 ago
Agileengine
México
About the Role
We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. This role blends platform engineering with incident command, using Terraform, CI/CD pipelines, and CSPM tools like Wiz. You will lead major-incident calls, own remediation follow-through, and build the playbooks that guide response.
What you will do
- Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP).
- Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations.
- Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment.
- Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads.
- Serve as Incident Commander on major and critical incidents — running the bridge, directing technical workstreams, making time-critical decisions, and coordinating cross-functional responders under pressure.
- Own the post-incident loop — track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups.
- Draft and send clear, accurate, audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle.
- Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution.
Must haves
- 5+ years of experience.
- In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles.
- Strong practical experience with Kubernetes, Terraform,
CI/CD orchestration, and Python/Go scripting.
- Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment.
- Proven track record of remediation follow-up — coordinating with teams and holding owners accountable until issues are fully closed.
- Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences.
- Direct experience authoring divisional/group incident-management playbooks and escalation procedures.
- Fully autonomous.
- Drives the architecture of complex automated runbooks and mentors Middle-level SREs.
- Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz.
- Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2).
- Upper-intermediate English level.
Nice to haves
- PagerDuty — hands-on experience with on-call scheduling, alert routing, and incident orchestration.
- ServiceNow — familiarity with incident, problem, and change management workflows and reporting.
Perks and Benefits
- Professional growth
Accelerate your professional journey with mentorship, TechTalks, and personalized growth roadmaps
- Competitive compensation
We match your ever-growing skills, talent, and contributions with competitive USD-based compensation and budgets for education, fitness, and team activities
- A selection of exciting projects
Join projects with modern solutions development and top-tier clients that include Fortune 500 enterprises and leading product brands
- Flextime
Tailor your schedule for an optimal work-life balance, by having the options of working from home and going to the office – whatever makes you the happiest and most productive.
📌 Site Reliability Engineer (México)
🏢 Agileengine
📍 México