08 oct
|
Epam Systems
|
México
08 oct
Epam Systems
México
We are seeking a
Chief Site Reliability Engineer
to strengthen reliability and DevOps maturity across mission-critical platforms in a fast-changing environment. You will drive resilient cloud and Kubernetes operations, improve CI/CD and release processes, and lead rapid incident response.
Lead reliability strategy for critical infrastructure to enable rapid business change Design resilient cloud architectures and operational patterns across AWS and Azure Build and improve CI/CD pipelines, source control practices, and release workflows Automate infrastructure and operational tasks using Python to reduce toil and risk Harden platform foundations across networking, compute, security, IAM, and configuration automation Operate and evolve Kubernetes usage patterns to improve stability and delivery speed Coordinate on-call response and resolve business-critical incidents under time pressure Investigate systemic issues, perform root-cause analysis, and drive corrective actions Partner with engineering teams to deliver high-leverage solutions over quick fixes Define and track reliability metrics and operational controls to measure maturity Extensive site reliability engineering experience (7+ years)
supporting critical infrastructure Strong cloud platform experience (7+ years) with leading providers, including Amazon Web Services and Microsoft Azure Proven leadership skills to set direction, influence stakeholders, and raise engineering standards Enterprise-scale release management experience delivering reliable software delivery processes Deep CI/CD expertise across pipelines, source control, and infrastructure automation Advanced Python programming skills to build tooling and automation Solid Kubernetes experience as a developer working with clusters and workloads Hands-on DevSecOps platform experience with GitLab preferred Strong analytical skills for complex problem-solving and strategic decision-making Upper-Intermediate English proficiency (B2, Upper-Intermediate) Amazon Web Services certification or equivalent hands-on expertise Microsoft Azure certification or equivalent hands-on expertise AI Architecture experience applied to reliability and operational decision-making AI Solution Engineering experience supporting platform automation and operations Gen AI Solutions Development exposure for operational tooling or incident workflows
📌 Chief Site Reliability Engineer (México)
🏢 Epam Systems
📍 México