Site Reliability Engineer (Ciudad de México)

Site Reliability Engineer (Ciudad de México)

21 ago
|
Infojini
|
Ciudad de México

21 ago

Infojini

Ciudad de México

Role: Site Reliability Engineering Location: LATAM Remote
Hire Type: Contract
Working Hours: 5 PM to 1 AM EST (Mon – Fri)

About the Role & Team
We are expanding our Site Reliability Engineering (SRE) organization. This role is part of newly established offshore SRE teams that will work in close partnership with our US-based engineering teams to ensure the reliability, availability, and performance of critical production systems.
This is a high-impact, front-line operations role focused on real-time incident response, proactive prevention, and continuous automation. Every minute matters—our SREs act decisively to prevent service degradation and protect the customer experience.

What You’ll Do
Incident Response & Command
· Act as the first responder to alerts and production incidents, rapidly assessing severity and initiating mitigation actions
· Serve as Incident Commander during major incidents, leading bridge calls with clarity and urgency
· Drive root cause isolation within 30 minutes for critical incidents whenever possible
· Communicate effectively across engineering, product, and leadership during high-pressure situations
· Maintain a strong presence on incident bridges—this role requires confidence, ownership, and clear decision-making

Proactive Reliability Engineering
· Identify patterns, trends, and signals to prevent incidents before they occur
· Continuously improve alert quality, reduce noise, and increase signal fidelity
· Partner with engineering teams to enhance system resilience and reliability

Automation & Toil Reduction

· Eliminate manual work by automating operational tasks, ticket handling, and repetitive workflows
· Build and improve tooling across incident response, observability, and operations
· Leverage AI-assisted development tools (e.g., Cursor, Claude) where they provide clear value

Platform & Systems Support

Troubleshoot across a hybrid ecosystem including:

· On-prem VMs (Linux & Windows; VMware)




· Cloud platforms (AWS, GCP, Azure)
· Containerized environments (Kubernetes clusters)

Diagnose and resolve issues across:
· Networking (connectivity, latency, DB access interruptions)
· Kubernetes (ingress, environment variables, cluster-level issues)
· CDN and traffic management layers (Akamai, waiting rooms – plus)

Required Technical Skills & Experience

Core Engineering & Operations
· Strong experience in incident management and triage in production environments
· Proven ability to troubleshoot complex distributed systems under pressure
· Solid understanding of Linux systems administration (including performance, networking, NTP, etc.)

Cloud & Infrastructure
· Hands-on experience with AWS core services (S3, Lambda, Load Balancers, ECS, EC2)
· Familiarity with GCP and/or Azure environments
· Experience operating in multi-cloud and hybrid environments

Containers & Orchestration
· Experience troubleshooting Kubernetes clusters (pods, ingress, configuration issues)
· Understanding of containerized application architectures

DevOps & CI/CD
Strong knowledge of DevOps practices and CI/CD pipelines
Hands-on experience with:
· Harness
· GitHub and/or GitLab

Application & Technology Stack Awareness

Working knowledge of:
· Java, Node.js, React-based applications
Understanding of database connectivity and dependencies across:
· Oracle, MariaDB, MSSQL (no DBA ownership, but strong troubleshooting awareness required)

Networking
Strong foundational knowledge of:
· TCP/IP, DNS,
· Load balancing and network troubleshooting
· Diagnosing connectivity issues between services and databases

Preferred Qualifications
Experience in large-scale enterprise (Fortune 500) environments supporting mission-critical applications
Prior experience as an Incident Commander or similar leadership role during outages
Familiarity with Akamai CDN and traffic management tools
Experience in high-volume, high-availability production environments

📌 Site Reliability Engineer (Ciudad de México)
🏢 Infojini
📍 Ciudad de México

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (ciudad de méxico) / ciudad de méxico

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (ciudad de méxico) / ciudad de méxico