06 ago
|
Epam Systems
|
Morelia
06 ago
Epam Systems
Morelia
EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential. We are seeking a Lead Site Reliability Engineer (SRE) to keep our cloud platforms reliable, scalable, and safe through software-driven operations and automation. You will reduce toil, strengthen availability, and protect production for teams delivering solutions across Financial Services, Insurance, and Retail. Join us to improve resilience and delivery speed while keeping downtime low—
Responsibilities Design, build and maintain cloud infrastructure using modern IaC practices such as Terraform or CloudFormation
Create and optimize CI/CD pipelines to automate software deployments, configuration management and repetitive operational tasks
Implement robust logging, monitoring and alerting systems to establish clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
Respond to production incidents and lead troubleshooting efforts to restore services
Run blameless post-mortems to identify root causes and prevent recurrence
Collaborate with software developers to optimize system performance and plan capacity
Ensure services can scale to handle growth and traffic spikes
Requirements Proven experience of 5+ years in systems administration, DevOps, or systems-oriented software development
Hands-on proficiency in at least one scripting or programming language such as Python, Bash, Go or Rust
Solid experience with public cloud providers such as AWS, Azure or GCP and containerization tools such as Docker and Kubernetes
Deep understanding of Linux/Unix administration and networking fundamentals such as TCP/IP, DNS and HTTP/SSL/TLS
Working knowledge of monitoring and observability tools such as Prometheus, Grafana or Datadog
Clear passion for automation, eliminating toil, and building resilient systems that fail gracefully
English proficiency at B2 (Upper-Intermediate) level or higher
We offer International projects with top brands
Work with global teams of highly skilled, diverse peers
Healthcare benefits
Employee financial programs
Paid time off and sick leave
Upskilling, reskilling and certification courses
Unlimited access to the LinkedIn Learning library and 22,000+ courses
General career opportunities
Volunteer and community involvement opportunities
EPAM Employee Groups
Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
#J-18808-Ljbffr
📌 Lead Site Reliability Engineer (SRE) (Morelia)
🏢 Epam Systems
📍 Morelia