04 sep
|
Infovision
|
Ciudad de México
04 sep
Infovision
Ciudad de México
About the Role We are looking for a Lead Site Reliability Engineer to own the reliability, scalability, and automation strategy across our production systems. This is a hands-on leadership role for someone who thrives at the intersection of software engineering and operations — someone who doesn't just fix outages, but engineers them out of existence. You will drive operational excellence across cloud infrastructure, CI/CD pipelines, observability, and compliance, while mentoring engineers and setting the technical bar for reliability practices across the organization. What You Will Do Lead reliability engineering initiatives across distributed, cloud-based systems, balancing feature velocity with system stability Design and implement scalable, resilient infrastructure across multi-cloud environments (AWS, GCP, Azure, PCF, VMware) Own incident response, root cause analysis, and post-mortem culture; drive down MTTR through automation Build and maintain robust CI/CD pipelines and deployment automation (Jenkins, XLR, Ansible, Chef, Habitat) Champion observability best practices using tools such as Splunk, Dynatrace, App Dynamics, and Thousand Eyes to catch issues before they impact users Drive an automation-first culture, reducing operational toil through scripting and tooling (Python, Go, Shell, Power Shell, Groovy) Ensure operational governance and compliance aligned with ITIL frameworks Manage certificate lifecycle and network security posture (Venafi, ECMS, Open SSL, F5, load balancers)
Mentor and provide technical guidance to a team of SRE and Dev Ops engineers, acting as a force multiplier for engineering best practices Partner cross-functionally with development, security, and infrastructure teams to embed reliability into the software lifecycle What You Will Bring 8+ years of experience in Site Reliability Engineering, Dev Ops, or Infrastructure Engineering, with demonstrated technical leadership Strong programming and scripting proficiency in Java, Python, Go, Shell, Power Shell, and/or Groovy, with a solid grounding in algorithms and data structures Hands-on experience designing and operating scalable systems across AWS, GCP, Azure, and PCF/VMware environments Proficiency with database technologies including Postgres, Cassandra, and Redis, along with strong SQL skills Deep experience with observability and monitoring stacks: Splunk, Dynatrace, App Dynamics, Blaze Meter, Thousand Eyes, Browser Stack Solid understanding of CI/CD toolchains and branching strategies (SCM, Artifactory, Sonar Qube, Jenkins, XLR, Ansible, Chef Infra, Habitat) Working knowledge of networking fundamentals: TCP/IP, load balancers (F5), proxies, FTP/SFTP, Wireshark Experience with certificate management and PKI (Venafi, ECMS, Open SSL) Familiarity with the ITIL framework and operational governance practices Proven ability to lead through influence, mentor engineers, and drive reliability culture at scale
📌 Lead sre engineer (Ciudad de México)
🏢 Infovision
📍 Ciudad de México