04 sep
|
Infovision
|
México
about the role
we are looking for a lead site reliability engineer to own the reliability, scalability, and automation strategy across our production systems. This is a hands-on leadership role for someone who thrives at the intersection of software engineering and operations — someone who doesn't just fix outages, but engineers them out of existence.
you will drive operational excellence across cloud infrastructure, ci/cd pipelines, observability, and compliance, while mentoring engineers and setting the technical bar for reliability practices across the organization.
what you will do
- lead reliability engineering initiatives across distributed, cloud-based systems, balancing feature velocity with system stability
- design and implement scalable, resilient infrastructure across multi-cloud environments (aws, gcp, azure, pcf, vmware)
- own incident response, root cause analysis, and post-mortem culture; drive down mttr through automation
- build and maintain robust ci/cd pipelines and deployment automation (jenkins, xlr, ansible, chef, habitat)
- champion observability best practices using tools such as splunk, dynatrace, appdynamics, and thousandeyes to catch issues before they impact users
- drive an automation-first culture, reducing operational toil through scripting and tooling (python, go, shell, powershell, groovy)
- ensure operational governance and compliance aligned with itil frameworks
- manage certificate lifecycle and network security posture (venafi, ecms, openssl, f5, load balancers)
- mentor and provide technical guidance to a team of sre and devops engineers, acting as a force multiplier for engineering best practices
- partner cross-functionally with development, security, and infrastructure teams to embed reliability into the software lifecycle
what you will bring
- 8+ years of experience in site reliability engineering, devops, or infrastructure engineering, with demonstrated technical leadership
- strong programming and scripting proficiency in java, python, go, shell, powershell, and/or groovy, with a solid grounding in algorithms and data structures
- hands-on experience designing and operating scalable systems across aws, gcp, azure, and pcf/vmware environments
- proficiency with database technologies including postgres, cassandra, and redis, along with strong sql skills
- deep experience with observability and monitoring stacks: splunk, dynatrace, appdynamics, blazemeter, thousandeyes, browserstack
- solid understanding of ci/cd toolchains and branching strategies (scm, artifactory, sonarqube, jenkins, xlr, ansible, chef infra, habitat)
- working knowledge of networking fundamentals: tcp/ip, load balancers (f5), proxies, ftp/sftp, wireshark
- experience with certificate management and pki (venafi, ecms, openssl)
- familiarity with the itil framework and operational governance practices
- proven ability to lead through influence, mentor engineers, and drive reliability culture at scale
📌 Lead sre engineer (México)
🏢 Infovision
📍 México