We are looking for a hands-on Senior Site Reliability Engineer to help maintain, enhance, and support a Java services ecosystem in close collaboration with an SRE peer and a backend engineering team. You will strengthen reliability, observability, and operational readiness while participating in on-call support.ResponsibilitiesProvide on-call support for Java backend identity services during business hoursTroubleshoot complex production issues using logs and telemetry and drive root-cause resolutionPrepare and deploy patches to address issues in cloud infrastructureImprove service reliability by implementing practical changes that reduce errors and instabilityBuild and refine metrics and dashboards to surface platform health and service behaviorMonitor SLOs and propose code changes that improve SLO attainment as issues ariseCreate and improve runbooks to standardize operational response and reduce time to recoveryCommunicate incidents and operational risks clearly in writing during live responseCollaborate closely with engineers to align operational practices with service ownershipRequirements3+ years of Site Reliability Engineering or DevOps experience supporting distributed systemsStrong on-call support experience for production services and incident response during business hoursProven experience with Amazon Web Services in production environmentsHands-on experience with Amazon DynamoDB and Amazon ElastiCacheStrong Git skills for collaborating on operational and reliability code changesSolid Gradle knowledge for building and maintaining Java-based servicesStrong troubleshooting skills using logs and telemetry to identify root causesClear written communication skills for documenting and reporting operational issues during incidentsProactive learning mindset to absorb complex information quickly and apply it under pressureUpper-Intermediate English proficiency (B2)Nice to haveKubernetesTerraformGrafanaNew RelicApache KafkaWe offerInternational projects with top brandsWork with general teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedInEPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
#J-18808-Ljbffr
📌 Senior Site Reliability Engineer (Ciudad de México)
🏢 Epam Systems
📍 Ciudad de México
Postulate a este anuncio
Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.