Site Reliability Engineer/ 100% Remote in Mexico (México)

Site Reliability Engineer/ 100% Remote in Mexico (México)

21 ago
|
Pyramid Consulting
|
México

21 ago

Pyramid Consulting

México

Job Profile: Site Reliability Engineer

Job Type: Long-Time based contract Job opportunity

Location: 100% Remote in Mexico

Job Description:

We are looking for an experienced Senior Site Reliability Engineer (SRE) / Reliability Engineer to improve the stability, availability, and reliability of production systems. This role focuses on the intersection of systems engineering, data science, and resilience engineering , using incident data and operational insights to identify trends, uncover root causes, and develop actionable reliability strategies.

Key Responsibilities

- Analyze production incidents, system performance data, and incident logs to identify reliability trends and opportunities for improvement.
- Develop tooling and processes to transform raw incident data into actionable reliability strategies .
- Design and implement solutions that improve system stability, availability, scalability, and operational resilience.
- Build and maintain observability solutions covering metrics, structured logging, alerting, and distributed tracing .
- Investigate complex production issues and perform root cause analysis across distributed systems.
- Analyze operational data using SQL and statistical techniques to identify anomalies, trends, and performance patterns.
- Apply statistical methods such as outlier detection and regression analysis to system performance and reliability data.
- Develop automation and engineering tools using Golang, Java, Python, C++, or other programming languages .
- Design, deploy, and maintain containerized workloads using Kubernetes .
- Manage and troubleshoot cloud infrastructure, with strong experience in Google Cloud Platform (GCP) .
- Partner with engineering, DevOps, data, and product teams to establish reliability best practices and improve operational processes.
- Apply Resilience Engineering principles to understand how systems, processes, and human decision-making influence reliability.




- Identify opportunities to reduce recurring incidents, operational toil, and mean time to recovery.
- Translate complex technical failures and reliability issues into clear business-impact reports for technical and non-technical stakeholders.
- Contribute to a culture of continuous learning, incident analysis, and proactive reliability improvement.

Required Skills & Experience

- 4+ years of experience in SRE, DevOps, Systems Engineering, or related roles managing production environments at scale.
- Strong hands-on experience with SQL and data analysis .
- Strong programming experience in one or more of Golang, Java, Python, or C++ .
- Deep understanding of observability , including alerting systems, distributed tracing, structured logging, and metrics collection.
- Experience with Kubernetes and container orchestration.
- Strong experience with GCP cloud infrastructure .
- Experience analyzing system performance and operational data using statistical methods.
- Knowledge of outlier detection, regression analysis, trend analysis, and anomaly detection .
- Strong understanding of distributed systems, production troubleshooting, and reliability engineering.
- Ability to perform effective root cause analysis and develop preventative reliability strategies.
- Strong understanding of Resilience Engineering and the human factors that influence system reliability.
- Excellent communication skills, with the ability to translate technical incidents into clear business impact and executive-level reporting .

Preferred Skills

- Experience with Prometheus, Grafana, OpenTelemetry, Jaeger, Datadog, Splunk, or similar observability platforms .
- Experience with Python-based data analysis and statistical libraries.
- Experience building reliability dashboards and operational analytics.
- Knowledge of SLOs, SLIs, SLAs, error budgets, and incident management .
- Experience with incident response, postmortems, and reliability improvement programs.
- Familiarity with Resilience Engineering, Chaos Engineering, and distributed systems reliability .

📌 Site Reliability Engineer/ 100% Remote in Mexico (México)
🏢 Pyramid Consulting
📍 México

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer/ 100% remote in mexico (méxico) / méxico

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer/ 100% remote in mexico (méxico) / méxico