19 ago
|
Pyramid Consulting
|
Chihuahua
19 ago
Pyramid Consulting
Chihuahua
Job Profile: Site Reliability Engineer
Job Type: Long-Time based contract Job opportunity
Location: 100% Remote in Mexico
Job Description:
We are looking for an experienced Senior Site Reliability Engineer (SRE) / Reliability Engineer to improve the stability, availability, and reliability of production systems. This role focuses on the intersection of systems engineering, data science, and resilience engineering, using incident data and operational insights to identify trends, uncover root causes, and develop actionable reliability strategies.
Key Responsibilities
- Analyze production incidents, system performance data, and incident logs to identify reliability trends and opportunities for improvement.
- Develop tooling and processes to transform raw incident data into actionable reliability strategies.
- Design and implement solutions that improve system stability, availability, scalability, and operational resilience.
- Build and maintain observability solutions covering metrics, structured logging, alerting, and distributed tracing.
- Investigate complex production issues and perform root cause analysis across distributed systems.
- Analyze operational data using SQL and statistical techniques to identify anomalies, trends, and performance patterns.
- Apply statistical methods such as outlier detection and regression analysis to system performance and reliability data.
- Develop automation and engineering tools using Golang, Java, Python, C++, or other programming languages.
- Design, deploy, and maintain containerized workloads using Kubernetes.
- Manage and troubleshoot cloud infrastructure, with strong experience in Google Cloud Platform (GCP).
- Partner with engineering, Dev Ops, data, and product teams to establish reliability best practices and improve operational processes.
- Apply Resilience Engineering principles to understand how systems, processes, and human decision-making influence reliability.
- Identify opportunities to reduce recurring incidents, operational toil, and mean time to recovery.
- Translate complex technical failures and reliability issues into clear business-impact reports for technical and non-technical stakeholders.
- Contribute to a culture of continuous learning, incident analysis, and proactive reliability improvement.
Required Skills & Experience
- 4+ years of experience in SRE, Dev Ops, Systems Engineering, or related roles managing production environments at scale.
- Strong hands-on experience with SQL and data analysis.
- Strong programming experience in one or more of Golang, Java, Python, or C++.
- Deep understanding of observability, including alerting systems, distributed tracing, structured logging, and metrics collection.
- Experience with Kubernetes and container orchestration.
- Strong experience with GCP cloud infrastructure.
- Experience analyzing system performance and operational data using statistical methods.
- Knowledge of outlier detection, regression analysis, trend analysis, and anomaly detection.
- Strong understanding of distributed systems, production troubleshooting, and reliability engineering.
- Ability to perform effective root cause analysis and develop preventative reliability strategies.
- Strong understanding of Resilience Engineering and the human factors that influence system reliability.
- Excellent communication skills, with the ability to translate technical incidents into clear business impact and executive-level reporting.
Preferred Skills
- Experience with Prometheus, Grafana, Open Telemetry, Jaeger, Datadog, Splunk, or similar observability platforms.
- Experience with Python-based data analysis and statistical libraries.
- Experience building reliability dashboards and operational analytics.
- Knowledge of SLOs, SLIs, SLAs, error budgets, and incident management.
- Experience with incident response, postmortems, and reliability improvement programs.
- Familiarity with Resilience Engineering, Chaos Engineering, and distributed systems reliability.
📌 Site Reliability Engineer/ 100% Remote in Mexico (Chihuahua)
🏢 Pyramid Consulting
📍 Chihuahua