Principal Site Reliability Engineer (Estado de México)

Principal Site Reliability Engineer (Estado de México)

12 ago
|
Oracle
|
Estado de México

12 ago

Oracle

Estado de México

Job Description

As a Principal member of the Site Reliability Engineering (SRE) team, you'll take ownership of highly available systems, influence service design, and work across teams to drive resiliency, automation, and operational excellence. This is a hands‑on engineering role where deep infrastructure knowledge meets software engineering expertise, adecuado for experienced SREs ready to take the lead.

This is not a fully remote role but a hybrid role. Requires in‑office presence at least 3 days a week in Guadalajara.

Responsibilities
- Lead the design, automation, and support of OCI services with a focus on resiliency, security, scalability, and performance.
- Own and improve the end‑to‑end reliability metrics (SLOs, SLAs, KPIs) for your services.
- Design and implement high‑availability architectures and standards for large‑scale distributed systems.
- Serve as the ultimate escalation point for complex operational issues, using a deep understanding of service topologies and interdependencies.




- Architect and build automation and orchestration tools that reduce manual work and prevent problem recurrence.
- Collaborate with development teams to improve service designs, optimize deployments, and implement best practices for operational efficiency.
- Guide technical decision‑making and mentor junior SREs and developers across teams.
- Participate in and lead post‑mortems, root‑cause analysis, and preventative design changes.
- Contribute to capacity planning, demand forecasting, and long‑term service scalability strategies.
- Participate in a rotational on‑call schedule to ensure the health and availability of production services.

What We’re Looking For
- Advanced experience with Linux systems administration.
- Strong programming skills in Python (with automation libraries).
- Advanced Bash/Shell scripting.
- Deep understanding of distributed systems, networking, and service architecture.
- Solid knowledge of databases and how they behave in production (SQL or No

📌 Principal Site Reliability Engineer (Estado de México)
🏢 Oracle
📍 Estado de México

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: principal site reliability engineer (estado de méxico) / estado de méxico

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: principal site reliability engineer (estado de méxico) / estado de méxico