31 jul
|
Epam Systems
|
México
31 jul
Epam Systems
México
Join our team as a **Lead Site Reliability Engineer** dedicated to providing advanced support for critical Azure-based systems.
**Responsibilities**
- Resolve complex incidents to ensure system availability
- Maintain reliability and performance of Azure-based enterprise infrastructure
- Deploy observability, monitoring, and logging tools
- Automate infrastructure management with Terraform and scripting technologies
- Improve system performance and uptime through centralized monitoring
- Collaborate with multiple teams to enhance service reliability
- Perform root cause analysis and oversee postmortems for incidents
- Configure deployment pipelines in Azure DevOps for secure workflows
- Write and maintain automation scripts for incident recovery and recurring tasks
- Enhance monitoring frameworks with platforms like Prometheus and Grafana
- Respond promptly to incidents to meet SLA expectations
- Facilitate integration of monitoring data from Azure and AWS environments
- Advance service reliability and observability practices continuously
- Document processes and incident resolutions thoroughly
- Take part in Agile team events and balance task priorities
**Requirements**:
- Minimum 5 years’ expertise in site reliability engineering or comparable DevOps roles
- 1+ years of demonstrated leadership experience
- Knowledge of Azure services, including AKS, Azure Monitor, Application Insights, Log Analytics, Cosmos DB, and PostgreSQL
- Expertise in infrastructure automation using Azure DevOps and Terraform
- Proficiency in scripting languages such as Bash, PowerShell, and Python
- Skills in monitoring tools including Prometheus and Grafana
- Background in incident management and ITSM processes with analytical capability for root cause investigations
- Competency in resolving technical challenges promptly in high-pressure situations
- Experience in Agile workflows and fast-paced operational environments
- Flexibility to communicate effectively in written and verbal formats for teamwork and documentation
- Capability to configure alerts that prevent SLA breaches proactively
- Understanding of cloud scaling techniques and security best practices
- Knowledge of Kubernetes administration for orchestration tasks
- Ability to collaborate with diverse functional teams seamlessly
- English proficiency of B2 or higher
**Nice to have**
- Background in AWS services, such as EKS, RDS, CloudWatch, and X-Ray
- Familiarity with distributed logging systems and tools for incident automation
- Certifications such as Microsoft Azure Administrator or AWS Certified DevOps Engineer
- Understanding of Kubernetes configurations for scaling and advanced networking setups
- Proficiency in observability tools such as OpenSearch for AWS environments
**We offer**
- Career plan and real growth opportunities
- Unlimited access to LinkedIn learning solutions
- International Mobility Plan within 25 countries
- Constant training, mentoring,
online corporate courses, eLearning and more
- English classes with a certified teacher
- Support for employee’s initiatives (Algorithms club, toastmasters, agile club and more)
- Enjoyable working environment (Gaming room, napping area, amenities, events, sport teams and more)
- Flexible work schedule and dress code
- Collaborate in a multicultural environment and share best practices from around the globe
- Hired directly by EPAM & 100% under payroll
- Law benefits (IMSS, INFONAVIT, 25% vacation bonus)
- Major medical expenses insurance: Life, Major medical expenses with dental & visual coverage (for the employee and direct family members)
- 13 % employee savings fund, capped to the law limit
- Grocery coupons
- 30 days December bonus
- Employee Stock Purchase Plan
- 12 vacations days plus 4 floating days
- Official Mexican holidays, plus 5 extra holidays (Maundry Thursday and Friday, November 2nd, December 24th & 31st)
- Monthly non-taxable amount for the electricity and internet bills
EPAM is a leading general provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.
📌 Lead Site Reliability Engineer (México)
🏢 Epam Systems
📍 México