04 sep
|
Infovision
|
Ciudad de México
04 sep
Infovision
Ciudad de México
About the Role We are looking for a Lead Site Reliability Engineer to own the reliability, scalability, and automation strategy across our production systems. This is a hands-on leadership role for someone who thrives at the intersection of software engineering and operations — someone who doesn't just fix outages, but engineers them out of existence.
You will drive operational excellence across cloud infrastructure, CI/CD pipelines, observability, and compliance, while mentoring engineers and setting the technical bar for reliability practices across the organization.
What You Will Do
Lead reliability engineering initiatives across distributed, cloud-based systems, balancing feature velocity with system stability
Design and implement scalable, resilient infrastructure across multi-cloud environments (AWS, GCP, Azure, PCF, VMware)
Own incident response, root cause analysis, and post-mortem culture; drive down MTTR through automation
Build and maintain robust CI/CD pipelines and deployment automation (Jenkins, XLR, Ansible, Chef, Habitat)
Champion observability best practices using tools such as Splunk, Dynatrace, AppDynamics, and ThousandEyes to catch issues before they impact users
Drive an automation-first culture, reducing operational toil through scripting and tooling (Python, Go, Shell, PowerShell, Groovy)
Ensure operational governance and compliance aligned with ITIL frameworks
Manage certificate lifecycle and network security posture (Venafi, ECMS, OpenSSL, F5, load balancers)
Mentor and provide technical guidance to a team of SRE and DevOps engineers, acting as a force multiplier for engineering best practices
Partner cross-functionally with development, security, and infrastructure teams to embed reliability into the software lifecycle
What You Will Bring
~8+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering, with demonstrated technical leadership
~ Strong programming and scripting proficiency in Java, Python, Go, Shell, PowerShell, and/or Groovy, with a solid grounding in algorithms and data structures
~ Hands-on experience designing and operating scalable systems across AWS, GCP, Azure, and PCF/VMware environments
~ Proficiency with database technologies including Postgres, Cassandra, and Redis, along with strong SQL skills
~ Deep experience with observability and monitoring stacks: Splunk, Dynatrace, AppDynamics, BlazeMeter, ThousandEyes, BrowserStack
~ Solid understanding of CI/CD toolchains and branching strategies (SCM, Artifactory, SonarQube, Jenkins, XLR, Ansible, Chef Infra, Habitat)
~ Working knowledge of networking fundamentals: TCP/IP, load balancers (F5), proxies, FTP/SFTP, Wireshark
~ Experience with certificate management and PKI (Venafi, ECMS, OpenSSL)
~ Familiarity with the ITIL framework and operational governance practices
~ Proven ability to lead through influence, mentor engineers, and drive reliability culture at scale
📌 LIDER TECNICO_CLOUD AZURE (Ciudad de México)
🏢 Infovision
📍 Ciudad de México