Are you passionate about building robust observability ecosystems and scaling reliability for mission-critical healthcare platforms? We are looking for a Senior DevOps/SRE Engineer to join our team at Luxoft!
In this role, you will lead the charge on client-scoped SLO monitoring, automated dashboard generation across 110–140 services, and building internal applications that give our clients clear SLA visibility. Plus, if you love leveraging cutting-edge tech—including AI-assisted development tools like Claude Code and GitHub Copilot—this is the exact environment for you.
? Why Luxoft?
At Luxoft, we thrive on solving complex problems through product engineering, hypothesis testing, and empowering our people to drive innovation. Certified as one of the Top 10 companies to work for in Mexico by the Great Place to Work Institute, we offer:
- True flexibility to balance your work and lifestyle.
- High-impact projects in dynamic Agile environments.
- A results-driven culture focused on continuous growth, training, and robust career support.
?️ What You Will Do
- Design and maintain robust, scalable SLO-based monitoring and alerting solutions.
- Create and optimize PromQL queries and multi-window burn-rate alerts.
- Build and manage Grafana dashboards using configuration-as-code practices.
- Develop automation tools and monitoring configuration generators in Python or TypeScript.
- Contribute to Terraform-based observability infrastructure.
- Validate monitoring signals and continuously improve alert quality across distributed systems.
- Collaborate closely with engineering and platform teams to maximize reliability and operational visibility.
? What We Are Looking For
- Experience: Strong professional background in Observability, SRE, or Platform Engineering within cloud-native/distributed systems environments.
- Query Expertise: Advanced PromQL (or equivalent) skills, including the critical ability to spot and correct misleading query results.
- SLO Mastery: Hands-on experience designing and implementing SLOs and multi-window burn-rate alerting.
- Code as Config: Proven experience with Grafana provisioning and configuration-as-code.
- Automation: Solid capability to build custom automation tools and code generators using Python or TypeScript.
- AI-Driven Workflow: Hands-on experience with AI-assisted development tools (e.g., Claude Code, GitHub Copilot). AI-assisted engineering is an integrated and expected part of our development workflow here!
- Analytical Mindset: A healthy skepticism toward telemetry data, with a commitment to validating signals through multiple independent sources before operationalizing alerts.