03 sep
|
Ntt Data
|
Guadalajara
03 sep
Ntt Data
Guadalajara
NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now. We are currently seeking a Site Reliability Engineering (SRE) / Lead Engineer to join our team in Guadalajara, Jalisco (MX-JAL), Mexico (MX). Site Reliability Engineer (SRE) / Lead Engineer candidate will have deep expertise in Application Performance Monitoring (APM), Infrastructure as Code (Ia C), automation, and distributed tracing using Open Telemetry.
As a SRE lead, he will guide the design, implementation, and continuous improvement of observability solutions, ensuring system reliability, performance, and scalability while fostering best practices in SRE and Dev Ops.
Responsibilities: u00b7 Lead the strategic development and management of observability and reliability frameworks across the organization, ensuring alignment with business goals and technical requirements.
u00b7 Design and implementation of monitoring and observability solutions, collaborating with engineering teams to define standards and best practices.
u00b7 Manage Infrastructure as Code (Ia C) initiatives using Terraform, coordinating with cloud and infrastructure teams to ensure scalable and secure deployments.
u00b7 Drive automation strategies for monitoring, alerting, and logging pipelines, focusing on process improvements and operational efficiency.
u00b7 Develop and maintain comprehensive observability roadmaps, including distributed tracing, logging, and metrics collection strategies.
u00b7 Collaborate with product management, sales, and pre-sales teams to provide technical expertise and support during solution design and customer engagements.
u00b7 Lead cross-functional teams to enhance CI/CD pipelines and deployment reliability, ensuring smooth integration of observability tools and practices.
u00b7 Engage with vendors and strategic partners to evaluate, select, and integrate observability and monitoring solutions, ensuring alignment with organizational needs and fostering strong collaborative relationships.
u00b7 Mentor and develop junior engineers and analysts, fostering a culture of reliability, observability, and operational excellence.
Qualifications
u00b7 8-10+ years of experience in SRE, Observability, or Dev Ops roles, with leadership responsibilities.
u00b7 Hands-on experience with Open Telemetry for distributed tracing and observability instrumentation.
u00b7 Proven expertise with Application Performance Monitoring (APM) tools such as New Relic, Datadog, App Dynamics, or Dynatrace.
u00b7 Strong proficiency in Infrastructure as Code (Ia C) using Terraform.
u00b7 Solid understanding of cloud platforms including AWS, GCP, or Azure.
u00b7 Experience with automation/configuration management tools like Ansible, Chef, or Puppet.
u00b7 Deep knowledge of CI/CD pipelines and tools such as Git Hub Actions, Jenkins, or Azure Dev Ops.
u00b7 Experience managing Kubernetes and containerized environments (Docker, Helm).
u00b7 Familiarity with log aggregation and analysis platforms like ELK Stack or Splunk.
📌 Site reliability engineering (sre) / lead engineer (Guadalajara)
🏢 Ntt Data
📍 Guadalajara