09 ago
|
Epam Systems
|
México
09 ago
Epam Systems
México
EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture.
Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential. We are seeking a skilled Lead Compute Platform SRE to support EPAM's Compute Managed Services project for our client.
The role focuses on KTLO (Keep the Lights On) activities, ensuring 24x7 monitoring, incident management, and operational stability across multi-cloud environments (GCP, AWS, Azure). The SRE will drive observability improvements, automate processes, and maintain compliance while collaborating with cross-functional teams to deliver high-quality compute services.
Responsibilities Perform continuous 24x7 monitoring of compute platforms using tools such as ELK and PagerDuty
Manage incidents and problems across servers, middleware, operating systems, and cloud platforms, including troubleshooting, root cause analysis (RCA), and resolution
Execute repaving activities, change management processes, and disaster recovery procedures
Ensure security and vulnerability compliance, including user management and certificate lifecycle management
Handle service requests, configuration updates, and audit-related data extracts
Develop and maintain Standard Operating Procedures (SOPs) for infrastructure operations
Collaborate on cell-based automation improvements and drive continuous service enhancements Requirements A minimum of 5 years of relevant experience
At least one year of experience leading and managing teams
Experience working with cloud platforms such as GCP, AWS, and Azure
Proficiency in OS administration across Windows and Linux environments
Proficiency in automation tools such as Ansible, Terraform, Python, and Bash
Strong knowledge of observability tools such as ELK and Grafana
Solid understanding of incident management processes
Experience using GitHub for version control and collaborative development
Experience in security hardening, vulnerability management, and compliance practices
Excellent problem-solving, communication, and collaboration skills
Familiarity with disaster recovery and operational recovery processes
English level B2 or higher, with strong written and verbal communication skills We offer International projects with top brands
Work with integral teams of highly skilled, diverse peers
Healthcare benefits
Employee financial programs
Paid time off and sick leave
Upskilling, reskilling and certification courses
Unlimited access to the LinkedIn Learning library and 22,000+ courses
Global career opportunities
Volunteer and community involvement opportunities
EPAM Employee Groups
Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
📌 Lead Site Reliability Engineer (México)
🏢 Epam Systems
📍 México