What you'll do reliability & operations - own availability, latency, and scalability across saa s and ai systems - define and enforce slos, slis, and error budgets - participate in a general on-call rotation (~1 week every 4 weeks) - lead incident response and drive blameless postmortems with systemic fixes platform & infrastructure - architect and operate on-premise and multi-region, multi-cloud environments - manage large-scale kubernetes workloads - build and evolve infrastructure using terraform and ansible - improve system resilience, fault isolation, and capacity planning ai/ml & automation - build and scale agentic ai systems for triage, anomaly detection, and self-healing - ensure reliability of model serving infrastructure - operate, optimize and scale distributed systems what you bring - 5+ years in sre , production engineering,
or platform engineering - strong experience with cloud providers (aws/gcp/oci), kubernetes, and ia c (terraform/ansible) - proficiency in python, go, or type script - experience with distributed systems and ai/ml platforms - deep understanding of slos, observability, and incident management - strong bias toward automation and system-level problem solving culture & growth - blameless, transparent postmortems - current mix: ~70% operations / 30% engineering , with active investment in automation and ai-driven toil reduction - clear growth path into staff/principal technical leadership or management
📌 Senior site reliability engineer/devops (Ciudad de México)
🏢 RCS TECH
📍 Ciudad de México
Postulate a este anuncio
Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.