Senior Site Reliability Engineer (México)

Senior Site Reliability Engineer (México)

31 jul
|
Epam Systems
|
México

31 jul

Epam Systems

México

Join our team as a **Senior Site Reliability Engineer** focused on delivering advanced support for critical Azure-based systems.
**Responsibilities**
- Troubleshoot and resolve complex incidents to maintain system uptime
- Ensure reliability and performance of Azure-based enterprise infrastructure
- Implement observability, monitoring, and logging solutions
- Automate infrastructure provisioning and deployment using Terraform and scripting
- Optimize system performance and uptime through proactive monitoring and alerting
- Collaborate with cross-functional teams to improve service reliability
- Conduct root cause analysis and postmortems for incident management
- Manage deployment pipelines in Azure DevOps for secure and scalable workflows
- Develop and maintain automation scripts for routine tasks and incident recovery
- Enhance monitoring frameworks with tools like Prometheus and Grafana
- React quickly to incidents to avoid SLA degradation
- Integrate monitoring data from Azure and AWS environments
- Support continuous improvement of service reliability and observability practices
- Document technical processes and incident reports
- Participate in Agile team activities and prioritize competing tasks
**Requirements**:
- Minimum 3 years of experience in site reliability engineering or related DevOps roles
- Hands-on experience with Azure services, including AKS, Azure Monitor, Application Insights, Log Analytics, Cosmos DB, and PostgreSQL
- Strong expertise in Azure DevOps and Terraform for infrastructure automation
- Proficient scripting skills in Bash, PowerShell, and Python




- Experience with monitoring and observability tools such as Prometheus and Grafana
- Solid background in incident management and ITSM processes with root cause analysis capabilities
- Ability to troubleshoot and debug complex technical issues in real-time
- Experience working in fast-paced Agile environments
- Strong verbal and written communication skills for collaboration and reporting
- Proactive approach to setting alerts and preventing SLA degradation
- Experience with cloud infrastructure scaling and security best practices
- Knowledge of Kubernetes administration and orchestration
- Ability to collaborate effectively with cross-functional teams
- English language proficiency at B2 level or above
**Nice to have**
- Hands-on experience with AWS services including EKS, RDS, CloudWatch, and X-Ray
- Familiarity with distributed logging pipelines and incident automation tools
- Knowledge of advanced Kubernetes use cases for scaling and network configurations
- Certifications such as Microsoft Azure Administrator or AWS Certified DevOps Engineer
- Experience with observability tools like OpenSearch for AWS workloads
**We offer**
- Career plan and real growth opportunities
- Unlimited access to LinkedIn learning solutions
- International Mobility Plan within 25 countries
- Constant training, mentoring, online corporate courses,



eLearning and more
- English classes with a certified teacher
- Support for employee’s initiatives (Algorithms club, toastmasters, agile club and more)
- Enjoyable working environment (Gaming room, napping area, amenities, events, sport teams and more)
- Flexible work schedule and dress code
- Collaborate in a multicultural environment and share best practices from around the globe
- Hired directly by EPAM & 100% under payroll
- Law benefits (IMSS, INFONAVIT, 25% vacation bonus)
- Major medical expenses insurance: Life, Major medical expenses with dental & visual coverage (for the employee and direct family members)
- 13 % employee savings fund, capped to the law limit
- Grocery coupons
- 30 days December bonus
- Employee Stock Purchase Plan
- 12 vacations days plus 4 floating days
- Official Mexican holidays, plus 5 extra holidays (Maundry Thursday and Friday, November 2nd, December 24th & 31st)
- Monthly non-taxable amount for the electricity and internet bills
EPAM is a leading integral provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.

📌 Senior Site Reliability Engineer (México)
🏢 Epam Systems
📍 México

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: senior site reliability engineer (méxico) / méxico

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: senior site reliability engineer (méxico) / méxico