We are looking for a
Lead Site Reliability Engineer
to strengthen critical infrastructure reliability and accelerate safe delivery of change in a fast-moving environment. You will elevate DevOps maturity across tooling, processes, and engineering practices while solving complex reliability challenges.
Own reliability outcomes for critical infrastructure and set SRE engineering standards Design and implement scalable DevOps tooling and processes that improve delivery speed and safety Build and maintain CI/CD workflows and source control practices using GitLab where appropriate Automate infrastructure operations with Python to reduce toil and improve consistency Improve Kubernetes-based workflows to support dependable deployments and runtime stability Diagnose and resolve business-critical incidents during on-call shifts with urgency and rigor Evaluate and harden infrastructure domains such as networking, compute, security, IAM,
and configuration automation Drive enterprise-grade release management practices across environments Partner with stakeholders to translate reliability needs into actionable engineering plans Promote long-term engineering solutions over quick fixes through root-cause analysis and follow-up actions 5+ years of site reliability engineering experience in production environments 5+ years of cloud platform experience with a leading provider Proven leadership skills to drive reliability improvements and mentor engineers Enterprise-scale release management experience across complex systems Strong CI/CD knowledge covering pipelines, source control, and automation practices Advanced Python programming skills for automation and engineering solutions Hands-on Kubernetes experience as a developer in delivery workflows Strong infrastructure fundamentals across networking, compute, security, and IAM Excellent analytical skills for strategic thinking and complex problem solving Effective incident response skills, including on-call ownership for critical issues English proficiency: B2 Upper-Intermediate Amazon Web Services expertise in production environments Microsoft Azure expertise in production environments AI Architecture experience for reliability-aware AI platform design AI Solution Engineering experience supporting clients with scalable implementations Gen AI Solutions Development experience in enterprise settings
📌 Lead Site Reliability Engineer (México)
🏢 Epam Systems
📍 México
Postulate a este anuncio
Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.