Site Reliability Engineer (México)

Site Reliability Engineer (México)

19 ago
|
Infojini
|
México

19 ago

Infojini

México

Role: Site Reliability Engineering

Location: LATAM Remote

Hire Type: Contract

Working Hours: 5 PM to 1 AM EST (Mon – Fri)

About the Role & Team

We are expanding our Site Reliability Engineering (SRE) organization. This role is part of newly established offshore SRE teams that will work in close partnership with our US-based engineering teams to ensure the reliability, availability, and performance of critical production systems.

This is a high-impact, front-line operations role focused on real-time incident response, proactive prevention, and continuous automation. Every minute matters—our SREs act decisively to prevent service degradation and protect the customer experience.

What You’ll Do

Incident Response & Command

· Act as the first responder to alerts and production incidents, rapidly assessing severity and initiating mitigation actions

· Serve as Incident Commander during major incidents, leading bridge calls with clarity and urgency

· Drive root cause isolation within 30 minutes for critical incidents whenever possible

· Communicate effectively across engineering, product, and leadership during high-pressure situations

· Maintain a strong presence on incident bridges—this role requires confidence, ownership, and clear decision-making

Proactive Reliability Engineering

· Identify patterns, trends, and signals to prevent incidents before they occur

· Continuously improve alert quality, reduce noise, and increase signal fidelity

· Partner with engineering teams to enhance system resilience and reliability

Automation & Toil Reduction

· Eliminate manual work by automating operational tasks, ticket handling, and repetitive workflows

· Build and improve tooling across incident response, observability, and operations

· Leverage AI-assisted development tools (e.g., Cursor, Claude) where they provide clear value

Platform & Systems Support

Troubleshoot across a hybrid ecosystem including:

· On-prem VMs (Linux & Windows; VMware)





· Cloud platforms (AWS, GCP, Azure)

· Containerized environments (Kubernetes clusters)

Diagnose and resolve issues across:

· Networking (connectivity, latency, DB access interruptions)

· Kubernetes (ingress, environment variables, cluster-level issues)

· CDN and traffic management layers (Akamai, waiting rooms – plus)

Required Technical Skills & Experience

Core Engineering & Operations

· Strong experience in incident management and triage in production environments

· Proven ability to troubleshoot complex distributed systems under pressure

· Solid understanding of Linux systems administration (including performance, networking, NTP, etc.)

Cloud & Infrastructure

· Hands-on experience with AWS core services (S3, Lambda, Load Balancers, ECS, EC2)

· Familiarity with GCP and/or Azure environments

· Experience operating in multi-cloud and hybrid environments

Containers & Orchestration

· Experience troubleshooting Kubernetes clusters (pods, ingress, configuration issues)

· Understanding of containerized application architectures

Dev Ops & CI/CD

Strong knowledge of Dev Ops practices and CI/CD pipelines

Hands-on experience with:

· Harness

· Git Hub and/or Git Lab

Application & Technology Stack Awareness

Working knowledge of:

· Java, Node.js, React-based applications

Understanding of database connectivity and dependencies across:

· Oracle, MariaDB, MSSQL (no DBA ownership, but strong troubleshooting awareness required)

Networking

Strong foundational knowledge of:

· TCP/IP, DNS, HTTP(S)

· Load balancing and network troubleshooting

· Diagnosing connectivity issues between services and databases

Preferred Qualifications

Experience in large-scale enterprise (Fortune 500) environments supporting mission-critical applications

Prior experience as an Incident Commander or similar leadership role during outages

Familiarity with Akamai CDN and traffic management tools

Experience in high-volume, high-availability production environments

📌 Site Reliability Engineer (México)
🏢 Infojini
📍 México

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (méxico) / méxico

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (méxico) / méxico