21 ago
|
Infojini
|
Ciudad de México
21 ago
Infojini
Ciudad de México
Role: Site Reliability Engineering Location: LATAM Remote Hire Type: Contract Working Hours: 5 PM to 1 AM EST (Mon – Fri) About the Role & Team We are expanding our Site Reliability Engineering (SRE) organization. This role is part of newly established offshore SRE teams that will work in close partnership with our US-based engineering teams to ensure the reliability, availability, and performance of critical production systems. This is a high-impact, front-line operations role focused on real-time incident response, proactive prevention, and continuous automation. Every minute matters—our SREs act decisively to prevent service degradation and protect the customer experience. What You’ll Do Incident Response & Command · Act as the first responder to alerts and production incidents, rapidly assessing severity and initiating mitigation actions · Serve as Incident Commander during major incidents, leading bridge calls with clarity and urgency · Drive root cause isolation within 30 minutes for critical incidents whenever possible · Communicate effectively across engineering, product, and leadership during high-pressure situations · Maintain a strong presence on incident bridges—this role requires confidence, ownership, and clear decision-making Proactive Reliability Engineering · Identify patterns, trends, and signals to prevent incidents before they occur · Continuously improve alert quality, reduce noise, and increase signal fidelity · Partner with engineering teams to enhance system resilience and reliability Automation & Toil Reduction · Eliminate manual work by automating operational tasks, ticket handling, and repetitive workflows · Build and improve tooling across incident response, observability, and operations · Leverage AI-assisted development tools (e.g., Cursor, Claude) where they provide clear value Platform & Systems Support Troubleshoot across a hybrid ecosystem including:
· On-prem VMs (Linux & Windows; VMware) · Cloud platforms (AWS, GCP, Azure) · Containerized environments (Kubernetes clusters) Diagnose and resolve issues across: · Networking (connectivity, latency, DB access interruptions) · Kubernetes (ingress, environment variables, cluster-level issues) · CDN and traffic management layers (Akamai, waiting rooms – plus) Required Technical Skills & Experience Core Engineering & Operations · Strong experience in incident management and triage in production environments · Proven ability to troubleshoot complex distributed systems under pressure · Solid understanding of Linux systems administration (including performance, networking, NTP, etc.) Cloud & Infrastructure · Hands-on experience with AWS core services (S3, Lambda, Load Balancers, ECS, EC2) · Familiarity with GCP and/or Azure environments · Experience operating in multi-cloud and hybrid environments Containers & Orchestration · Experience troubleshooting Kubernetes clusters (pods, ingress, configuration issues) · Understanding of containerized application architectures Dev Ops & CI/CD Strong knowledge of Dev Ops practices and CI/CD pipelines Hands-on experience with: · Harness · Git Hub and/or Git Lab Application & Technology Stack Awareness Working knowledge of: · Java, Node.js, React-based applications Understanding of database connectivity and dependencies across: · Oracle, Maria DB, MSSQL (no DBA ownership, but strong troubleshooting awareness required) Networking Strong foundational knowledge of: · TCP/IP, DNS, · Load balancing and network troubleshooting · Diagnosing connectivity issues between services and databases Preferred Qualifications Experience in large-scale enterprise (Fortune 500) environments supporting mission-critical applications Prior experience as an Incident Commander or similar leadership role during outages Familiarity with Akamai CDN and traffic management tools Experience in high-volume, high-availability production environments
📌 Site reliability engineer (Ciudad de México)
🏢 Infojini
📍 Ciudad de México