Are you ready to harness observability to keep world-scale science running at pace? Could you lead the strategy that helps researchers trust the platforms behind computational chemistry, genomics and AI-driven discovery while raising your own profile through visible, high-impact work?
As a Senior DevOps Engineer focused on Observability, you will shape how our Scientific Computing Platform is monitored and understood, turning signals into insights that prevent incidents and accelerate research. You will evolve our current Prometheus and Grafana landscape and guide a multi-year shift toward an enterprise-grade observability capability using commercial tooling such as Datadog, Dynatrace, New Relic, or Grafana Cloud — ensuring we can scale confidently as demand grows.
This hands-on senior role sits within a collaborative, globally distributed platform engineering team operating across cloud-native services and large-scale bare-metal estates. You will combine automation, reliability engineering,
and deep telemetry expertise so scientists can compute with confidence and get treatments to patients faster.
Accountabilities
- Observability Platform Ownership:Operate, improve, and scale the Prometheus and Grafana stack on Kubernetes; define alerting strategies and build actionable dashboards that cut time-to-detection and time-to-recovery for critical services.
- Migration and Future Tooling:Lead our contribution to the centralized enterprise observability initiative; evaluate and prepare for adoption of commercial platforms (Datadog, Dynatrace, New Relic, Grafana Cloud), building a clear path from today/'s stack to tomorrow/'s shared capability.
- Automation and GitOps:Increase reliability and repeatability through GitOps workflows and CI/CD pipelines; automate deployments and configuration to reduce manual effort and eliminate drift across environments.
- Platform Administration:Administer Ansible, Vault, Consul, Prometheus, and Grafana to underpin