objective of the role
the staff infrastructure engineer – sre is a senior technical leader responsible for architecting and scaling the reliability, availability, and performance of critical infrastructure platforms. This role combines deep technical expertise, cross-team influence, and a strategic mindset to drive high-impact initiatives across engineering teams, ensuring operational excellence and long-term system resilience.
main responsibilities
• design and evolve complex infrastructure systems to ensure scalability, reliability, and security of platform services.
• define long-term architectural vision and influence infrastructure roadmap across multiple teams and services.
• lead strategic sre initiatives, collaborating with cross-functional teams to improve system resiliency and reduce operational toil.
• identify systemic issues and lead root cause analysis efforts, establishing long-term corrective measures.
• champion best practices in observability, including metrics, logging, and alerting frameworks for production systems.
• design and implement advanced automation solutions to optimize infrastructure performance and operational workflows.
• partner with security and compliance teams to ensure infrastructure meets regulatory and organizational standards.
• mentor senior and mid-level engineers, fostering knowledge sharing and technical growth across the organization.
• serve as a technical advisor to leadership, providing insight on system health, risk, and architectural tradeoffs.
• collaborate with product, engineering, and data teams to support the needs of autonomous squads while maintaining infrastructure cohesion.
• lead capacity planning efforts across critical services to anticipate future growth and ensure infrastructure scalability.
• own and continuously improve disaster recovery strategies and testing plans to maintain business continuity.
• contribute to and enforce standards for infrastructure documentation, design reviews, and change management processes.
• foster a culture of reliability engineering, automation-first mindset, and operational accountability.
• embody and promote spin’s cultural values, acting as a role model of collaboration, innovation, and ownership.
• promote an autonomous work culture by encouraging self-management, accountability, and proactive problem-solving among team members.
• serve as a spin culture ambassador to foster and maintain a positive, inclusive, and dynamic work environment that aligns with the company's values and culture.
required knowledge and experience
• bachelor’s degree in computer science, information systems, or equivalent practical experience.
• 7+ years of experience in infrastructure, site reliability, or platform engineering roles.
• proven experience designing and operating reliable systems at scale, including distributed systems and cloud-native platforms.
• advanced knowledge of infrastructure-as-code tools, container orchestration systems (e.g., kubernetes), and ci/cd practices.
• strong programming and scripting skills (e.g., python, go, bash).
• deep understanding of monitoring, alerting, and incident management practices.
• experience working with cloud platforms (aws, gcp, or azure) and hybrid environments.
• demonstrated ability to drive cross-team collaboration and influence engineering practices across domains.
• strong leadership presence, excellent communication skills, and ability to present complex ideas to technical and non-technical audiences.
• high level of autonomy, initiative, and problem-solving skills in dynamic and fast-paced environments
📌 Staff sre sr (Xico)
🏢 Spin
📍 Xico