Site reliability engineer (México)

Site reliability engineer (México)

14 ago
|
Pyramid Consulting
|
México

14 ago

Pyramid Consulting

México

job description:

site reliability engineer (sre) / lead engineer candidate will have deep expertise in application performance monitoring (apm), infrastructure as code (iac), automation, and distributed tracing using opentelemetry. As a sre lead, he will guide the design, implementation, and continuous improvement of observability solutions, ensuring system reliability, performance, and scalability while fostering best practices in sre and devops.

responsibilities:

- lead the strategic development and management of observability and reliability frameworks across the organization, ensuring alignment with business goals and technical requirements.
- design and implementation of monitoring and observability solutions, collaborating with engineering teams to define standards and best practices.
- manage infrastructure as code (iac) initiatives using terraform, coordinating with cloud and infrastructure teams to ensure scalable and secure deployments.
- drive automation strategies for monitoring, alerting, and logging pipelines, focusing on process improvements and operational efficiency.
- develop and maintain comprehensive observability roadmaps, including distributed tracing, logging, and metrics collection strategies.
- collaborate with product management, sales, and pre-sales teams to provide technical expertise and support during solution design and customer engagements.




- lead cross-functional teams to enhance ci/cd pipelines and deployment reliability, ensuring smooth integration of observability tools and practices.
- engage with vendors and strategic partners to evaluate, select, and integrate observability and monitoring solutions, ensuring alignment with organizational needs and fostering strong collaborative relationships.
- mentor and develop junior engineers and analysts, fostering a culture of reliability, observability, and operational excellence.

qualifications

- 8-10+ years of experience in sre, observability, or devops roles, with leadership responsibilities.
- hands-on experience with opentelemetry for distributed tracing and observability instrumentation.
- proven expertise with application performance monitoring (apm) tools such as new relic, datadog, appdynamics, or dynatrace.
- strong proficiency in infrastructure as code (iac) using terraform.
- solid understanding of cloud platforms including aws, gcp, or azure.
- experience with automation/configuration management tools like ansible, chef, or puppet.
- deep knowledge of ci/cd pipelines and tools such as github actions, jenkins, or azure devops.
- experience managing kubernetes and containerized environments (docker, helm).
- familiarity with log aggregation and analysis platforms like elk stack or client.
- excellent leadership, communication, and collaboration skills.

📌 Site reliability engineer (México)
🏢 Pyramid Consulting
📍 México

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (méxico) / méxico

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (méxico) / méxico