08 oct
|
Epam Systems
|
México
08 oct
Epam Systems
México
We are seeking a
Senior Site Reliability Engineer
to strengthen critical infrastructure reliability and accelerate delivery through robust DevOps practices. You will drive improvements across CI/CD, cloud platforms, and operational readiness while solving high-impact production issues.
Design reliability strategies for critical infrastructure and services Build and maintain CI/CD pipelines and release workflows using GitLab Automate infrastructure operations and tooling with Python Operate cloud environments across AWS and Azure to meet availability goals Harden and standardize infrastructure practices across networking, security, IAM, and compute Coordinate incident response during on-call shifts and restore service quickly Implement Kubernetes-based deployment and operational patterns to improve stability Analyze reliability signals and root causes to prevent repeat incidents Improve DevOps processes and engineering capabilities to enable faster change Partner with stakeholders to prioritize resilience work over short-term fixes 3+ years of site reliability engineering experience in cloud environments Proven leadership ability to influence reliability practices across teams Enterprise-scale release management experience supporting frequent deployments Strong cloud platform knowledge across Amazon Web Services and Microsoft Azure Advanced Python programming skills for automation and tooling Solid Kubernetes skills using clusters as a developer Deep CI/CD and source control knowledge with GitLab or similar DevSecOps platforms Strong infrastructure fundamentals across networking, compute, security, IAM, and configuration automation Strong analytical skills for complex problem solving under pressure Upper-Intermediate English proficiency (B2) Reliable on-call readiness to assess and resolve business-critical issues Amazon Web Services expertise, including design patterns for resilient systems Microsoft Azure expertise, including governance and operational best practices AI Architecture experience applied to platform reliability and automation AI Solution Engineering experience for production-grade AI-enabled operations Gen AI Solutions Development experience focused on operational use cases
📌 Senior Site Reliability Engineer (México)
🏢 Epam Systems
📍 México