Lead SRE Engineer (Centro)

Lead SRE Engineer (Centro)

24 sep
|
CLOUDSUFI
|
Centro

24 sep

CLOUDSUFI

Centro

- Lead Site Reliability Engineering initiatives across enterprise production environments.
- Drive reliability, availability, scalability, and performance improvements.
- Lead P1/P2 incident response, RCA, and postmortems.
- Define and improve SLI, SLO, and Error Budget practices.
- Build and maintain observability, monitoring, dashboards, and alerts.
- Automate operational processes using Python/Bash and APIs.
- Manage Infrastructure as Code using Terraform.
- Support AWS cloud infrastructure, including Kubernetes/EKS, ECS/Fargate, Lambda, RDS/Aurora, SQS/SNS, and related services.
- Mentor SRE/DevOps engineers and collaborate with development, security, platform, and architecture teams.
- Contribute to AI-enabled reliability and AIOps initiatives.

Must-Have Experience





Observability / Monitoring – Mandatory

Hands-on experience with Datadog or an equivalent APM, logging, and monitoring platform such as New Relic, Dynatrace, Grafana/Prometheus, Splunk, ELK/OpenSearch, CloudWatch, or AppDynamics.

Experience should include:

- APM
- Metrics and logs
- Distributed tracing
- Dashboards and monitoring
- Alerting
- Production troubleshooting

Hands-on experience with PagerDuty or an equivalent alerting/paging platform , such as Opsgenie, ServiceNow, Splunk On-Call, xMatters, or Grafana Alerting.

#J-18808-Ljbffr

📌 Lead SRE Engineer (Centro)
🏢 CLOUDSUFI
📍 Centro

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: lead sre engineer (centro) / centro

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: lead sre engineer (centro) / centro