08 oct
|
Epam Systems
|
México
08 oct
Epam Systems
México
We are seeking a
Lead Site Reliability Engineer
to strengthen critical infrastructure reliability and accelerate DevOps maturity for high-impact services. You will design scalable automation, improve CI/CD and release practices, and lead rapid incident response.
Design reliability strategies and SRE practices for business-critical infrastructure Build automation and tooling in Python to improve stability, consistency, and operational leverage Develop and maintain CI/CD workflows and source control practices using GitLab Lead incident response during on-call rotations and restore service for business-critical issues Improve release management processes to support enterprise-scale delivery Harden cloud infrastructure across networking, compute, security, and IAM controls Implement configuration automation to reduce manual work and prevent drift Operate and troubleshoot Kubernetes-based workloads and developer-facing platform usage Partner with engineering stakeholders to prioritize reliability work and manage change safely Assess systemic risks and drive corrective actions to prevent recurring incidents 5+ years of site reliability engineering or DevOps experience in cloud environments Hands-on experience with a leading cloud provider,
with practical work across Amazon Web Services and Microsoft Azure Leadership ability to guide technical direction and take ownership of critical infrastructure outcomes Project delivery experience improving DevOps tools, processes, and engineering maturity at scale Deep CI/CD knowledge across pipelines, source control, and release management Strong Kubernetes skills with practical usage as a developer Advanced Python programming skills for automation and tooling Enterprise-scale release management experience supporting complex systems Solid infrastructure fundamentals across networking, compute, security, IAM, and configuration automation Strong analytical skills to diagnose complex issues and identify high-leverage solutions Effective incident response skills, including on-call ownership and rapid restoration of service Upper-Intermediate English proficiency (B2, Upper-Intermediate) Amazon Web Services certification or proven advanced AWS operational experience Microsoft Azure certification or proven advanced Azure operational experience AI Architecture experience for reliability-focused platform design AI Solution Engineering experience integrating AI-enabled capabilities into operations Gen AI Solutions Development experience for operational intelligence and automation use cases
📌 Lead Site Reliability Engineer (México)
🏢 Epam Systems
📍 México