Job Overview
We are seeking an experienced Sr.
Principal
Engineer, Site Reliability (SRE) to establish and scale the observability and reliability foundation for our integral multi-cloud SaaS platform serving thousands of customers worldwide. This role will define enterprise standards for logging, metrics, tracing, alerting, dashboards, SLIs/SLOs, and service ownership, then work hands-on with Engineering and Operations teams to put those standards into practice. The successful candidate will bring deep experience operating observability at scale, modernizing telemetry and tooling, reducing alert noise and tool sprawl, and using actionable signals to improve detection, MTTR, performance, and reliability.
Off hours support as needed Success Metrics
Customer Impact: Reduced MTTD/MTTR and improved customer experience through faster detection, diagnosis, and recovery
Observability Adoption: Measurable adoption of common logging, metrics, tracing, alerting, dashboarding, and service ownership standards across critical services
Reliability Engineering: Expanded use of meaningful SLIs, SLOs, and error budgets to drive service health and engineering priorities
Signal Quality: Reduction in noisy, duplicate, and non-actionable alerts while improving coverage of critical customer journeys and dependencies
Tooling & Cost Efficiency: Improved observability platform efficiency through governance, consolidation, telemetry optimization, and reduced tool sprawl
Cross-functional Adoption: Strong partnership with Product and Engineering teams that translates standards into measurable production adoption About Us
ICIMS is a leading enterprise hiring platform that combines the scale and reliability of enterprise software with the transformative power of AI. Thousands of organizations across more than 200 countries and territories trust ICIMS to find and hire the people who shape their future and drive their business forward. Powered by insights from billions of hiring interactions, continuous AI innovation, and a highly extensible platform, ICIMS helps organizations turn talent acquisition into a competitive advantage.
For more than 25 years, ICIMS has delivered end-to-end hiring solutions that improve recruiting efficiency, reduce costs and create exceptional candidate experiences. ICIMS helps solve one of the biggest challenges businesses face today: building a workforce that can adapt, scale, and perform in an increasingly competitive and unpredictable talent market. We uniquely do that by combining enterprise-grade hiring technology, AI-powered insights and automation, and connected talent experiences to help organizations improve hiring outcomes while driving measurable impact.
Responsibilities
Technical Leadership
Provide strategic technical direction for a team of 5+ SRE engineers across one or more geographic regions (US, Ireland, or India)
Own the technical strategy and roadmap for enterprise observability and reliability capabilities in partnership with SRE, Engineering, Cloud, and Product teams
Define reference architectures, engineering patterns, and standards that teams can consistently apply in production
Drive architecture reviews and technical decision-making for complex observability, reliability, scalability, and performance challenges
Provide hands-on technical mentorship and guidance, raising observability and SRE engineering practices across teams
Incident Management & Response
Participate in enterprise-wide incident management, ensuring rapid detection, response, restoration, and prevention of recurring issues
Improve incident detection and triage through actionable telemetry, service health views, dependency context, and well-designed alerting
Develop and maintain runbooks, emergency response procedures, and operational readiness practices for critical services
Lead root cause and post-incident reviews, ensuring clear documentation and implementation of durable corrective actions
Participate in 24/7 on-call and escalation procedures and serve as a senior technical leader with Engineering and Incident Management during critical incidents
Observability Strategy & Standards
Establish and evolve enterprise standards for logs, metrics, traces, alerting, dashboards, instrumentation, and service ownership
Champion OpenTelemetry-first, vendor-neutral instrumentation patterns with consistent context, correlation, naming, and metadata across services
Implement meaningful SLIs, SLOs, error budgets, and service health views that connect technical signals to customer impact
Drive practical adoption and governance of observability standards, measuring coverage, signal quality, and operational effectiveness across teams
Platform Reliability, Automation & Tooling
Design and operate scalable observability platforms and telemetry pipelines using technologies such as Grafana, Prometheus, Sumo Logic, New Relic, and cloud-native services
Lead observability platform modernization, migration, and consolidation while maintaining coverage and controlling ingestion, retention, cardinality, and overall tooling cost
Use infrastructure-as-code, automation, self-service patterns, and automated remediation to make reliability practices repeatable and reduce operational overhead
Monitor and optimize multi-cloud infrastructure and core services across AWS, Azure, and GCP for reliability, performance, capacity, and operational efficiency Qualifications
Bachelor’s degree in computer science, Engineering, Information Systems, or related technical field
Equivalent combination of education and experience will be considered
Cloud certifications (AWS, Azure, or Google Cloud) Technical Experience
8+ years in SRE, DevOps, Infrastructure Engineering, or Observability Engineering roles with 4+ years in senior technical positions
Proven hands-on experience designing, implementing,
and operating observability capabilities at scale in large enterprise SaaS or cloud production environments
Deep experience across logging, metrics, distributed tracing, alerting, dashboards, and OpenTelemetry, with platforms such as Grafana, Prometheus, Sumo Logic, and New Relic
Strong multi-cloud and cloud-native experience, including AWS, containers, Kubernetes/ECS, Linux, and distributed application architectures
Experience designing scalable telemetry pipelines and managing sampling, retention, cardinality, data quality, and cost tradeoffs in high-volume environments
Leadership & Communication
Proven track record creating technical standards and successfully driving them from architecture into consistent production adoption across engineering teams
Experience serving as a senior technical leader during critical incidents and complex cross-team reliability initiatives
Strong communication and influencing skills with engineers, architects, product leaders, and senior stakeholders
Demonstrated ability to mentor technical teams, build alignment across organizational boundaries, and lead through influence
SRE & Operations
Demonstrated success implementing SRE principles in large-scale production environments, including practical use of SLIs, SLOs, and error budgets
Strong background in incident management, root cause analysis, operational readiness, automation, and continuous reliability improvement
Experience with ITIL frameworks and tools and with establishing service-level expectations for enterprise SaaS products
Preferred
Experience leading large-scale observability platform migrations, consolidation initiatives, and telemetry cost/governance programs
Infrastructure-as-code expertise with Terraform or CloudFormation; authentication and identity management systems knowledge is a plus EEO Statement iCIMS is a place where everyone belongs. We celebrate diversity and are committed to creating an inclusive environment for all employees. Our approach helps us to build a winning team that represents a variety of backgrounds, perspectives, and abilities.
So, regardless of how your diversity expresses itself, you can find a home here at iCIMS. We prohibit discrimination and harassment of any kind based on race, color, religion, national origin, sex (including pregnancy), sexual orientation, gender identity, gender expression, age, veteran status, genetic information, disability, or other applicable legally protected characteristics. If you’d like to request an accommodation due to a disability, please contact us at
[email protected].
Compensation And Benefits
Competitive health and wellness benefits include medical insurance (employee and dependent family members), personal accident and group term life insurance, bonding and parental leave, lifestyle spending account reimbursements, wellness services offerings, sick and casual/emergency days, paid holidays, tuition reimbursement, retirals (PF - employer contribution) and gratuity.
Benefits and eligibility may vary by location, role, and tenure.
Learn more here: https://careers.icims.com/benefits
📌 Senior Site Reliability Engineer (México)
🏢 iCIMS
📍 México