02 ago
|
Pentangle Tech Services | P5 Group
|
Guadalajara
02 ago
Pentangle Tech Services | P5 Group
Guadalajara
this is a highly hands-on role focused on platform reliability, production support, and infrastructure. The adecuado candidate enjoys solving complex operational problems, automating repeatable work, improving system resilience, and supporting service health in a fast-moving environment.
responsibilities:
· own and support the reliability of voice assistant services and the underlying cloud infrastructure.
· participate in an on-call rotation to support production systems outside normal business hours.
· lead and participate in incident response, including triage, escalation, mitigation, and restoration of service.
· drive blameless postmortems and ensure corrective actions are tracked through completion.
· work to improve customer experiences by strengthening service availability, latency, stability, and resilience against agreed service-level expectations.
· design, implement, and maintain infrastructure as code using tools such as terraform and atlantis.
· manage and extend gitops and deployment workflows using argocd and related ci/cd tooling.
· support and improve cloud and container platforms across aws and azure.
· independently manage virtual servers, containers, and orchestration platforms such as kubernetes.
· build, refine, and operate automation that reduces toil and improves operational efficiency.
· utilize and extend existing observability and reliability capabilities, including monitoring, alerting, logging, and diagnostics.
· review system performance and capacity trends to identify bottlenecks, support forecasting, and improve scalability.
· troubleshoot complex infrastructure, networking, and application runtime issues across distributed systems.
· assist with disaster recovery planning, validation, and recovery readiness.
· contribute to performance tuning and resilience improvements across infrastructure and services.
· document operational procedures,
support runbooks, and engineering knowledge to improve team effectiveness.
· coach and support other engineers by sharing operational best practices and reliability engineering approaches.
required qualifications:
· 3+ years of production experience working as a site reliability engineer, devops engineer, infrastructure engineer, or software engineer
· experience working with atlantis, argocd, or similar infrastructure and deployment automation tools.
· strong experience with aws and azure.
· expertise in terraform to create, modify, and manage infrastructure configurations or iac templates
· expertise in containerization technologies (docker & kubernetes) to build, package, and deploy optimized container images
· proficiency in designing, implementing, and maintaining complex ci/cd pipelines that span with increasing complexity and integrating across multiple environments
· knowledge in cloud platforms (aws) to optimize cloud resource utilization and costs throughout product lifecycle
· expertise in version control systems to perform branching, merging, and resolving merge conflicts
· versed within infosec policies and procedures to adhere to security standards/regulations and identify gaps in security architecture
· experience in monitoring and analytics platforms to set up monitors, alerts, and diagnostic tools for proactive issue detection, root cause analysis, and performance optimization across distributed systems
· understanding of cloud billing and cost management tools to recognize total costs
· ability to learn and apply new technologies, programming practices, patterns, and methods
· organized and detail-oriented
· ability to develop healthy working relationships and collaborate with peers and leaders
· exhibits integrity and high standards in work quality
· excellent verbal and written communication skills
· values diversity and differences amongst individuals in interactions
📌 Site reliability engineer (Guadalajara)
🏢 Pentangle Tech Services | P5 Group
📍 Guadalajara