08 oct
|
Fulcrum Digital
|
Ciudad de México
08 oct
Fulcrum Digital
Ciudad de México
Job Description
Who We Are
Fulcrum Digital is a general AI-first enterprise transformation company with over 25 years of experience. We partner with enterprises across financial services, insurance, healthcare, retail, manufacturing, higher education, and logistics to move from AI experimentation to scalable business outcomes. Fulcrum Digital works with over 100 global clients, including Fortune 500 enterprises, combining deep industry expertise with capabilities in enterprise AI, digital engineering, cloud modernisation, platform integration, and generative AI.
The Role
As a Site Reliability Engineer, you will own the health, stability, and performance of a production environment supporting one of our client engagements. You will define how applications are monitored and supported, drive automation across deployment and operations, and work closely with development teams to reduce incidents and improve resiliency over time. This role suits someone with a strong service ownership mindset who enjoys connecting the dots across a complex technology stack and working with a global team across multiple time zones.
What You'll Do
• Plan, manage, and oversee all aspects of the production environment.
• Define strategies for application performance monitoring and optimization in production.
• Design, develop, and standardize monitoring and alerting mechanisms for supported applications.
• Respond to incidents, improve the platform based on feedback, and measure the reduction of incidents over time.
• Take a holistic approach to problem solving during production events, connecting the dots across the full technology stack to optimize mean time to recover (MTTR).
• Analyze ITSM activities for the platform and provide a feedback loop to development teams on operational gaps or resiliency concerns.
• Engage in and improve the whole lifecycle of services, from inception and design through deployment, operation, and refinement.
• Support services before they go live through system design consulting, capacity planning, and launch reviews.
• Maintain live services by measuring and monitoring availability, latency, and overall system health.
• Support code deployments into multiple lower environments, supporting current processes while automating wherever possible.
• Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead on Dev Ops automation and best practices.
• Scale systems sustainably through automation, pushing for changes that improve both reliability and velocity.
• Collaborate with a global team spread across tech hubs in multiple geographies and time zones, sharing knowledge and explaining processes and procedures to others.
Requirements
Requirements
• Strong hands-on experience with Linux.
• Experience with monitoring tools such as Splunk, Dynatrace, or equivalent.
• Working knowledge of ITIL/ITSM practices.
• Strong troubleshooting skills across complex, multi-layered platforms.
• Proficiency in SQL and PL/SQL.
• Experience with Jenkins and CI/CD pipelines.
• Scripting experience with Groovy, YAML, and Shell.
• Experience with Git and Bitbucket.
• Hands-on experience with Kubernetes and AWS.
• Proven experience in production support leadership, including runbook and support model creation.
• Experience defining monitoring and alerting strategies.
• Experience with disaster recovery and resiliency planning.
• Experience with deployment readiness validation and operational process design.
• Solid background in root cause analysis and problem management.
• A service ownership mindset with a focus on continuous improvement and toil reduction.
• Strong communication skills and the ability to share knowledge across distributed teams.
• Availability to participate in rotational on-call duties and occasional off-hours work.
Requirements
• Strong hands-on experience with Linux. • Experience with monitoring tools such as Splunk, Dynatrace, or equivalent. • Working knowledge of ITIL/ITSM practices. • Strong troubleshooting skills across complex, multi-layered platforms. • Proficiency in SQL and PL/SQL. • Experience with Jenkins and CI/CD pipelines. • Scripting experience with Groovy, YAML, and Shell. • Experience with Git and Bitbucket. • Hands-on experience with Kubernetes and AWS. • Proven experience in production support leadership, including runbook and support model creation. • Experience defining monitoring and alerting strategies. • Experience with disaster recovery and resiliency planning. • Experience with deployment readiness validation and operational process design. • Solid background in root cause analysis and problem management. • A service ownership mindset with a focus on continuous improvement and toil reduction. • Strong communication skills and the ability to share knowledge across distributed teams.
📌 SRE - Ciudad de México
🏢 Fulcrum Digital
📍 Ciudad de México