01 ago
|
Epam Systems
|
Morelia
01 ago
Epam Systems
Morelia
EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.
Site Reliability Engineer to strengthen the reliability, scalability, and safety of production environments. You will bridge software development and operations through automation, observability, and incident response—apply now to help reduce downtime and enable fast, safe releases.
Responsibilities Design and maintain cloud infrastructure using Infrastructure as Code practices
Build and optimize CI/CD pipelines to automate deployments and operational workflows
Implement logging, monitoring, and alerting to improve observability and reliability
Define and track Service Level Objectives and Service Level Indicators with clear reporting
Respond to production incidents and drive rapid service restoration
Lead blameless post-mortems to identify root causes and prevent recurrence
Partner with engineers to improve performance, scalability, and capacity planning
Automate repetitive operational tasks to reduce toil and operational risk
Harden production environments to improve resilience and safe change practices
Requirements 2+ years of experience in site reliability engineering, DevOps,
or systems administration
Hands-on experience with Infrastructure as Code using Terraform or CloudFormation
Hands-on experience building and improving CI/CD pipelines for automated deployments
Strong troubleshooting and incident response leadership skills in production environments
Solid project skills to coordinate reliability work with software development teams
Proficiency in scripting or programming with Python, Bash, Go, or Rust
Cloud platform experience with AWS, Azure, or GCP
Containerization experience with Docker and Kubernetes
Deep Linux/Unix administration knowledge and networking fundamentals (TCP/IP, DNS, HTTP, SSL/TLS)
Strong communication skills with a reliability mindset focused on automation and reducing toil
Advanced English proficiency (C1, Advanced)
Nice to have Experience with Prometheus, Grafana, or Datadog
We offer International projects with top brands
Work with integral teams of highly skilled, diverse peers
Healthcare benefits
Employee financial programs
Paid time off and sick leave
Upskilling, reskilling and certification courses
Unlimited access to the LinkedIn Learning library and 22,000+ courses
Global career opportunities
Volunteer and community involvement opportunities
EPAM Employee Groups
Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
#J-18808-Ljbffr
📌 Site Reliability Engineer (Morelia)
🏢 Epam Systems
📍 Morelia