05 ago
|
Pentangle Tech Services | P5 Group
|
Jalisco
05 ago
Pentangle Tech Services | P5 Group
Jalisco
Job Title: Site Reliability Engineer Location:Av.
Punto Sur 312, floor 4, int 159 , Los Gavilanes, *****, Tlajomulco de Zúñiga, Jalisco Duration: Long Term Our Toyota Connected Labs team is seeking a Site Reliability Engineer to create, maintain, support, and improve complex cloud operations for cutting edge microservice based platforms.
On this team you will solve complex problems and work alongside talented DevOps Engineers, Data Engineers, Software Engineers, Data Scientists, Agile Delivery Leads, and Technical Product Managers who have a passion for driving innovation in mobility, safety, and convenience services.
We love people who think big and like to get their hands dirty to help us build exciting initiatives.
This is a highly hands-on role focused on platform reliability, production support, and infrastructure.
The adecuado candidate enjoys solving complex operational problems, automating repeatable work, improving system resilience, and supporting service health in a fast-moving environment.
Responsibilities: · Own and support the reliability of Voice Assistant services and the underlying cloud infrastructure.
· Participate in an on-call rotation to support production systems outside normal business hours.
· Lead and participate in incident response, including triage, escalation, mitigation, and restoration of service.
· Drive blameless postmortems and ensure corrective actions are tracked through completion.
· Work to improve customer experiences by strengthening service availability, latency, stability, and resilience against agreed service-level expectations.
· Design, implement,
and maintain infrastructure as code using tools such as Terraform and Atlantis.
· Manage and extend GitOps and deployment workflows using ArgoCD and related CI/CD tooling.
· Support and improve cloud and container platforms across AWS and Azure.
· Independently manage virtual servers, containers, and orchestration platforms such as Kubernetes.
· Build, refine, and operate automation that reduces toil and improves operational efficiency.
· Utilize and extend existing observability and reliability capabilities, including monitoring, alerting, logging, and diagnostics.
· Review system performance and capacity trends to identify bottlenecks, support forecasting, and improve scalability.
· Troubleshoot complex infrastructure, networking, and application runtime issues across distributed systems.
· Assist with disaster recovery planning, validation, and recovery readiness.
· Contribute to performance tuning and resilience improvements across infrastructure and services.
· Document operational procedures, support runbooks, and engineering knowledge to improve team effectiveness.
· Coach and support other engineers by sharing operational best practices and reliability engineering approaches.
Required Qualifications: · 3+ years of production experience working as a Site Reliability Engineer, DevOps Engineer, Infrastructure Engineer, or Software Engineer · Experience working with Atlantis, ArgoCD, or similar infrastructure and deployment automation tools.
· Strong experience with AWS and Azure.
· Expertise in Terraform to create, modify, and manage infrastructure configurations or IaC templates · Expertise in containerization technologies (Docker
📌 Site Reliability Engineer (Jalisco)
🏢 Pentangle Tech Services | P5 Group
📍 Jalisco