Experienced Support Service Delivery Manager with 12+ years of expertise leading global Application Support, Production Operations, Site Reliability Engineering (SRE), and Managed Services organizations supporting mission‐critical enterprise applications.Proven track record of delivering high‐availability services, driving operational excellence, and managing geographically distributed support teams across 24x7 follow‐the‐sun environments.Demonstrated expertise in Azure Cloud Operations, Application Support, Incident & Problem Management, Service Delivery Governance, Observability, and Continuous Service Improvement.Adept at leading teams responsible for supporting customer‐facing web applications, APIs, cloud‐native platforms, databases, and integrations while ensuring SLA compliance, operational stability, and superior customer experience.Strong background in driving reliability engineering practices, proactive monitoring, automation, self‐healing operations, and root cause elimination initiatives.Experienced in partnering with engineering, product, infrastructure, and business stakeholders to improve platform performance, reduce operational risk, and optimize support costs.Core Responsibilities Service Delivery & Operational Leadership Lead end-to-end delivery of Azure‐based Application Support and Site Reliability Engineering services, ensuring consistent achievement of service level agreements, operational KPIs, and customer expectations.Oversee day‐to‐day support operations, resource planning, workload management, service governance, and operational reporting.Serve as the primary escalation point for critical production incidents and customer‐impacting outages.Lead major incident response efforts, coordinate cross‐functional resolution teams, manage stakeholder communications,
and ensure timely restoration of services.Drive root cause analysis and corrective action plans to prevent recurring issues.Azure Application Support & Cloud Operations Provide operational oversight for Azure‐hosted applications and cloud services including Azure App Services, Azure Container Apps, Azure Static Web Apps, Azure API Management (APIM), Azure Database for Postgre SQL, Azure Blob Storage, Azure Cache for Redis, and Azure AI services.Ensure platform stability, performance optimization, scalability, and operational readiness.Azure Cache for Redis Reliability Engineering & Observability Drive adoption of SRE best practices focused on service reliability, availability, resiliency, and performance.Establish and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and operational health metrics.Govern observability frameworks leveraging Azure Monitor, Application Insights, Log Analytics, and Kusto Query Language (KQL) to enable proactive issue detection and rapid troubleshooting.Champion continuous service improvement initiatives focused on reducing operational toil, improving service quality, and increasing team productivity.Drive automation of repetitive operational activities through scripting, workflow automation, self‐healing mechanisms,
and operational runbooks.Identify opportunities for AIOps and intelligent automation to enhance service delivery efficiency.Act as the primary operational liaison for customers, business stakeholders, and leadership teams.Conduct regular service reviews, operational governance meetings, and performance reporting sessions.Provide insights into service trends, operational risks, improvement opportunities, and strategic recommendations.Lead, mentor, and develop high‐performing Site Reliability Engineers and Application Support teams.Foster a culture of accountability, customer focus, technical excellence, collaboration, and continuous learning.Support hiring, onboarding, performance management, and career development initiatives.Monitoring & Observability Operational Dashboards & Alerting Performance Monitoring & Capacity Management Service Management Change & Release Management Root Cause Analysis (RCA) Knowledge Management Automation & Operational Excellence Power Shell Runbook Automation Self‐Healing Operations Operational Process Optimization Key Strengths Integral 24x7 Support Operations Leadership Production Support & Site Reliability Engineering Customer‐Focused Service Delivery Executive Stakeholder Communication SLA & KPI Governance Incident Command & Crisis Management Reliability & Performance Optimization Automation & AIOps Adoption Operational Transformation & Continuous Improvement Business Impact Consistently delivers measurable improvements in service reliability, operational efficiency, customer satisfaction, and support productivity through proactive operations management, automation, observability, and disciplined service governance.Proven ability to transform reactive support organizations into proactive, data‐driven, reliability‐focused operations teams that enable business growth while reducing operational risk and support costs.
#J-*****-Ljbffr
📌 Azure App Support & Sre Delivery Leader (Jalisco)
🏢 E-IT
📍 Jalisco