DevOps / Site Reliability Engineer (León)

DevOps / Site Reliability Engineer (León)

12 ago
|
Agileengine
|
León

12 ago

Agileengine

León

Job Description
AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.

WHY JOIN US
If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!

ABOUT THE ROLE
We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience and 24/7 stability for a multi-cloud enterprise security program, serving as Incident Commander on major and critical incidents while also owning IaC, CI/CD pipelines, and CSPM telemetry. You will drive major-incident calls, own post-incident remediation follow-through, draft stakeholder communications, and develop divisional incident-management playbooks alongside multi-cloud security guardrails using Terraform and Wiz. The role requires 5+ years of SRE experience with hands-on incident command in a 24/7 financial services environment.

WHAT YOU WILL DO
- Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP).
- Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations.
- Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment.
- Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads.
- Serve as Incident Commander on major and critical incidents — running the bridge, directing technical workstreams, making time-critical decisions, and coordinating cross-functional responders under pressure.
- Own the post-incident loop — track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups.
- Draft and send clear, accurate,



audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle.
- Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution.

MUST HAVES
- 5+ years of experience .
- In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles .
- Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting .
- Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment.
- Proven track record of remediation follow-up — coordinating with teams and holding owners accountable until issues are fully closed.
- Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences.
- Direct experience authoring divisional/group incident-management playbooks and escalation procedures.
- Fully autonomous.
- Drives the architecture of complex automated runbooks and mentors Middle-level SREs.
- Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz .
- Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2) .
- Upper-intermediate English level.

NICE TO HAVES
- PagerDuty — hands-on experience with on-call scheduling, alert routing, and incident orchestration.




- ServiceNow — familiarity with incident, problem, and change management workflows and reporting.

PERKS AND BENEFITS
- Growth without limits : build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget
- Competitive compensation : get recognition that reflects your skills and impact, with regular performance and compensation reviews
- Flexibility : work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm
- Meaningful, modern projects : build impactful products using modern technologies alongside integral teams and leading brands
- Collaborative culture : join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized
- Well-being & support : access local well-being programs and people-focused support tailored to your location

Requirements
5+ years of experience. In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles. Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting. Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment. Proven track record of remediation follow-up — coordinating with teams and holding owners accountable until issues are fully closed. Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences. Direct experience authoring divisional/group incident-management playbooks and escalation procedures Fully autonomous. Drives the architecture of complex automated runbooks and mentors Middle-level SREs. Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz. Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2).

📌 DevOps / Site Reliability Engineer (León)
🏢 Agileengine
📍 León

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: devops / site reliability engineer (león) / león

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: devops / site reliability engineer (león) / león