26 ago
|
Fulcrum Digital
|
Xico
26 ago
Fulcrum Digital
Xico
Fulcrum Digital is a general AI-first enterprise transformation company with over 25 years of experience.
We partner with enterprises across financial services, insurance, healthcare, retail, manufacturing, higher education, and logistics to move from AI experimentation to scalable business outcomes.
With over 100 global clients, including Fortune 500 enterprises, we combine deep industry expertise with capabilities in enterprise AI, digital engineering, cloud modernisation, platform integration, and generative AI.The RoleWe are looking for a Site Reliability Engineer (SRE), based in Mexico City on a hybrid schedule, to own and optimize the health of our production environment.
You will be at the center of platform reliability, working across the full service lifecycle from design through operation, while collaborating with a global team across multiple time zones.What You'll DoPlan, manage, and oversee all aspects of a production environmentDefine strategies for application performance monitoring and optimization in productionRespond to incidents, drive platform improvements, and measure incident reduction over timeSupport code deployment across multiple lower environments, with a strong focus on automationDesign, develop, and standardize monitoring and alerting mechanismsTake a holistic, cross-stack approach to problem-solving during production events to optimize mean time to recovery (MTTR)Own the full service lifecycle,
from inception and design through deployment, operation, and refinementAnalyze ITSM activity and provide feedback loops to development teams on operational gaps and resiliency concernsSupport pre-launch activities, including system design consulting, capacity planning, and launch reviewsSupport CI/CD pipelines through validation and operational gating, championing DevOps best practicesMonitor availability, latency, and overall system health to keep services running smoothlyDrive sustainable scaling through automation and reliability-focused system changesPerform root cause analysis and on-call support on a rotational basisCollaborate with a global team spread across multiple tech hubs and time zonesShare knowledge and mentor others on processes and proceduresOccasional off-hours work requiredRequirementsMust HaveSignificant strength in monitoring, including SOP creation — Splunk and Dynatrace experience is requiredStrong hands-on experience with LinuxShell scripting proficiencyStrong application troubleshooting skillsSolid understanding of ITIL / ITSM processesAlso RequiredSQL knowledgeWorking knowledge of Git / BitbucketRoot cause analysis and capacity planning experiencePreferred QualificationsCard payment knowledge (payment flows, switching, settlements, authorization flows)Experience with Prometheus/GrafanaCloud experience — AWS and/or Azure#J-*****-Ljbffr
📌 Sre (Engineering & Administration Background) (Xico)
🏢 Fulcrum Digital
📍 Xico