25 sep
|
CLOUDSUFI
|
Jalisco
Role: Lead SRE EngineerLocation: Guadalajara, Jalisco, MexicoExperience: 6–10 yearsKey ResponsibilitiesLead Site Reliability Engineering initiatives across enterprise production environments.Drive reliability, availability, scalability, and performance improvements.Lead P1/P2 incident response, RCA, and postmortems.Define and improve SLI, SLO, and Error Budget practices.Build and maintain observability, monitoring, dashboards, and alerts.Automate operational processes using Python/Bash and APIs.Manage Infrastructure as Code using Terraform.Support AWS cloud infrastructure, including Kubernetes/EKS, ECS/Fargate, Lambda, RDS/Aurora, SQS/SNS, and related services.Mentor SRE/DevOps engineers and collaborate with development, security, platform,
and architecture teams.Contribute to AI-enabled reliability and AIOps initiatives.Must-Have ExperienceObservability / Monitoring – MandatoryHands-on experience with Datadog or an equivalent APM, logging, and monitoring platform, such as New Relic, Dynatrace, Grafana/Prometheus, Splunk, ELK/OpenSearch, CloudWatch, or AppDynamics.Experience should include:APMMetrics and logsDistributed tracingDashboards and monitoringAlertingProduction troubleshootingAlerting / Incident Management – MandatoryHands-on experience with PagerDuty or an equivalent alerting/paging platform, such as Opsgenie, ServiceNow, Splunk On-Call, xMatters, or Grafana Alerting.Experience should include:
📌 Lead Sre Engineer (Jalisco)
🏢 CLOUDSUFI
📍 Jalisco