06 ago
|
Pentangle Tech Services | P5 Group
|
Guadalajara
06 ago
Pentangle Tech Services | P5 Group
Guadalajara
This is a highly hands-on role focused on platform reliability, production support, and infrastructure. The adecuado candidate enjoys solving complex operational problems, automating repeatable work, improving system resilience, and supporting service health in a fast-moving environment. Responsibilities:
- Own and support the reliability of Voice Assistant services and the underlying cloud infrastructure.
- Participate in an on-call rotation to support production systems outside normal business hours.
- Lead and participate in incident response, including triage, escalation, mitigation, and restoration of service.
- Drive blameless postmortems and ensure corrective actions are tracked through completion.
- Work to improve customer experiences by strengthening service availability, latency, stability, and resilience against agreed service-level expectations.
- Design, implement, and maintain infrastructure as code using tools such as Terraform and Atlantis.
- Manage and extend Git Ops and deployment workflows using ArgoCD and related CI/CD tooling.
- Support and improve cloud and container platforms across AWS and Azure.
- Independently manage virtual servers, containers, and orchestration platforms such as Kubernetes.
- Build, refine, and operate automation that reduces toil and improves operational efficiency.
- Utilize and extend existing observability and reliability capabilities, including monitoring, alerting, logging, and diagnostics.
- Review system performance and capacity trends to identify bottlenecks, support forecasting, and improve scalability.
- Troubleshoot complex infrastructure, networking, and application runtime issues across distributed systems.
- Assist with disaster recovery planning, validation, and recovery readiness.
- Contribute to performance tuning and resilience improvements across infrastructure and services.
- Document operational procedures, support runbooks,
and engineering knowledge to improve team effectiveness.
- Coach and support other engineers by sharing operational best practices and reliability engineering approaches.
Required Qualifications:
- 3+ years of production experience working as a Site Reliability Engineer, Dev Ops Engineer, Infrastructure Engineer, or Software Engineer
- Experience working with Atlantis, ArgoCD, or similar infrastructure and deployment automation tools.
- Strong experience with AWS and Azure.
- Expertise in Terraform to create, modify, and manage infrastructure configurations or IaC templates
- Expertise in containerization technologies (Docker & Kubernetes) to build, package, and deploy optimized container images
- Proficiency in designing, implementing, and maintaining complex CI/CD pipelines that span with increasing complexity and integrating across multiple environments
- Knowledge in cloud platforms (AWS) to optimize cloud resource utilization and costs throughout product lifecycle
- Expertise in version control systems to perform branching, merging, and resolving merge conflicts
- Versed within Info Sec policies and procedures to adhere to security standards/regulations and identify gaps in security architecture
- Experience in monitoring and analytics platforms to set up monitors, alerts, and diagnostic tools for proactive issue detection, root cause analysis, and performance optimization across distributed systems
- Understanding of cloud billing and cost management tools to recognize total costs
- Ability to learn and apply new technologies, programming practices, patterns, and methods
- Organized and detail-oriented
- Ability to develop healthy working relationships and collaborate with peers and leaders
- Exhibits integrity and high standards in work quality
- Excellent verbal and written communication skills
- Values diversity and differences amongst individuals in interactions
📌 Site Reliability Performance Engineer (Guadalajara)
🏢 Pentangle Tech Services | P5 Group
📍 Guadalajara