Objective of the Role
The Staff Infrastructure Engineer – SRE is a senior technical leader responsible for architecting and scaling the reliability, availability, and performance of critical infrastructure platforms. This role combines deep technical expertise, cross-team influence, and a strategic mindset to drive high-impact initiatives across engineering teams, ensuring operational excellence and long-term system resilience.
Main Responsibilities
• Design and evolve complex infrastructure systems to ensure scalability, reliability, and security of platform services.
• Define long-term architectural vision and influence infrastructure roadmap across multiple teams and services.
• Lead strategic SRE initiatives, collaborating with cross-functional teams to improve system resiliency and reduce operational toil.
• Identify systemic issues and lead root cause analysis efforts, establishing long-term corrective measures.
• Champion best practices in observability, including metrics, logging, and alerting frameworks for production systems.
• Design and implement advanced automation solutions to optimize infrastructure performance and operational workflows.
• Partner with security and compliance teams to ensure infrastructure meets regulatory and organizational standards.
• Mentor senior and mid-level engineers, fostering knowledge sharing and technical growth across the organization.
• Serve as a technical advisor to leadership, providing insight on system health, risk, and architectural tradeoffs.
• Collaborate with Product, Engineering, and Data teams to support the needs of autonomous squads while maintaining infrastructure cohesion.
• Lead capacity planning efforts across critical services to anticipate future growth and ensure infrastructure scalability.
• Own and continuously improve disaster recovery strategies and testing plans to maintain business continuity.
• Contribute to and enforce standards for infrastructure documentation, design reviews, and change man
📌 STAFF SRE SR (México)
🏢 Spin
📍 México