SRE - Ciudad de México

SRE - Ciudad de México

08 oct
|
Fulcrum Digital
|
Ciudad de México

08 oct

Fulcrum Digital

Ciudad de México

Job Description

Who We Are

Fulcrum Digital is a global AI-first enterprise transformation company with over 25 years of experience. We partner with enterprises across financial services, insurance, healthcare, retail, manufacturing, higher education, and logistics to move from AI experimentation to scalable business outcomes. Fulcrum Digital works with over 100 general clients, including Fortune 500 enterprises, combining deep industry expertise with capabilities in enterprise AI, digital engineering, cloud modernisation, platform integration, and generative AI.

The Role

As a Site Reliability Engineer, you will own the health, stability, and performance of a production environment supporting one of our client engagements. You will define how applications are monitored and supported, drive automation across deployment and operations, and work closely with development teams to reduce incidents and improve resiliency over time. This role suits someone with a strong service ownership mindset who enjoys connecting the dots across a complex technology stack and working with a global team across multiple time zones.

What You'll Do

• Plan, manage, and oversee all aspects of the production environment.

• Define strategies for application performance monitoring and optimization in production.

• Design, develop, and standardize monitoring and alerting mechanisms for supported applications.

• Respond to incidents, improve the platform based on feedback, and measure the reduction of incidents over time.

• Take a holistic approach to problem solving during production events, connecting the dots across the full technology stack to optimize mean time to recover (MTTR).

• Analyze ITSM activities for the platform and provide a feedback loop to development teams on operational gaps or resiliency concerns.





• Engage in and improve the whole lifecycle of services, from inception and design through deployment, operation, and refinement.

• Support services before they go live through system design consulting, capacity planning, and launch reviews.

• Maintain live services by measuring and monitoring availability, latency, and overall system health.

• Support code deployments into multiple lower environments, supporting current processes while automating wherever possible.

• Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead on Dev Ops automation and best practices.

• Scale systems sustainably through automation, pushing for changes that improve both reliability and velocity.

• Collaborate with a global team spread across tech hubs in multiple geographies and time zones, sharing knowledge and explaining processes and procedures to others.

Requirements

Requirements

• Strong hands-on experience with Linux.

• Experience with monitoring tools such as Splunk, Dynatrace, or equivalent.

• Working knowledge of ITIL/ITSM practices.

• Strong troubleshooting skills across complex, multi-layered platforms.

• Proficiency in SQL and PL/SQL.

• Experience with Jenkins and CI/CD pipelines.

• Scripting experience with Groovy, YAML, and Shell.

• Experience with Git and Bitbucket.





• Hands-on experience with Kubernetes and AWS.

• Proven experience in production support leadership, including runbook and support model creation.

• Experience defining monitoring and alerting strategies.

• Experience with disaster recovery and resiliency planning.

• Experience with deployment readiness validation and operational process design.

• Solid background in root cause analysis and problem management.

• A service ownership mindset with a focus on continuous improvement and toil reduction.

• Strong communication skills and the ability to share knowledge across distributed teams.

• Availability to participate in rotational on-call duties and occasional off-hours work.

Requirements
• Strong hands-on experience with Linux. • Experience with monitoring tools such as Splunk, Dynatrace, or equivalent. • Working knowledge of ITIL/ITSM practices. • Strong troubleshooting skills across complex, multi-layered platforms. • Proficiency in SQL and PL/SQL. • Experience with Jenkins and CI/CD pipelines. • Scripting experience with Groovy, YAML, and Shell. • Experience with Git and Bitbucket. • Hands-on experience with Kubernetes and AWS. • Proven experience in production support leadership, including runbook and support model creation. • Experience defining monitoring and alerting strategies. • Experience with disaster recovery and resiliency planning. • Experience with deployment readiness validation and operational process design. • Solid background in root cause analysis and problem management. • A service ownership mindset with a focus on continuous improvement and toil reduction. • Strong communication skills and the ability to share knowledge across distributed teams.

📌 SRE - Ciudad de México
🏢 Fulcrum Digital
📍 Ciudad de México

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: sre - ciudad de méxico / ciudad de méxico

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: sre - ciudad de méxico / ciudad de méxico