03 oct
|
Epam Systems
|
México
03 oct
Epam Systems
México
We are looking for a
Site Reliability Engineer
to strengthen reliability, observability, and platform operations across cloud and Kubernetes environments. In this role, you will improve service health through automation, infrastructure as code and CI/CD practices. Apply now to help keep critical systems stable and scalable!
Maintain service reliability by handling L2 operations and incident response Operate and administer Kubernetes clusters to ensure stability and performance Build and improve CI/CD workflows using Azure DevOps and Azure Pipelines Automate operational tasks using scripting to reduce manual effort Define and maintain infrastructure as code using Terraform and Ansible Implement and refine observability using MELT signals to detect and resolve issues faster Coordinate problem resolution by analyzing root causes and proposing corrective actions Support secure and resilient cloud operations on Microsoft Azure Document operational procedures and share knowledge to improve support readiness 2+ years of site reliability engineering or DevOps experience Kubernetes administration experience supporting production workloads Azure DevOps and Azure Pipelines experience delivering CI/CD workflows Infrastructure as Code expertise with Terraform and Ansible Proficiency in scripting languages for automation tasks Strong troubleshooting skills across metrics, events, logs, and traces (MELT) Strong understanding of observability concepts and tools Azure fundamentals knowledge with AZ-900 or AZ-104 certification (or higher) Good communication skills for cross-team incident coordination English proficiency: B1+ level or higher Argo CD administration or implementation experience Experience with Apache Cassandra cluster or timeseries/NoSQL operations Familiarity with Grafana and Elastic Cloud or Elastic Stack HashiCorp Vault experience Knowledge of programming languages, such as Python, Angular, or Go
📌 Site Reliability Engineer (México)
🏢 Epam Systems
📍 México