responsibilities
- design and evolve gcp cloud architecture, including networking, interconnects, iam, and high-availability topology, using terraform.
- build and own ci/cd pipelines for infrastructure planning, review, testing, policy enforcement, drift detection, and progressive rollout.
- develop self-service platform capabilities and golden paths that enable engineers to provision infrastructure without hand-offs.
- strengthen metrics, logging, tracing, and alerting through the prometheus, thanos, grafana, loki, tempo, and alertmanager observability stack.
- operate gke clusters, helm-packaged workloads, rabbitmq and ibm mq message brokers, and data stores in production.
- participate in follow-the-sun on-call from apac hours, triage alerts, lead incident debugging and escalation, and drive post-mortem actions.
- embed sre practices including slis, slos, error budgets, and capacity planning into infrastructure operations.
requirements
- at least 5 years of experience in devops, platform/infrastructure, or sre roles operating large-scale, highly available, high-performance production systems.
- deep hands-on experience designing google cloud platform architecture, including landing zones, networking, iam, and high-availability topology.
- strong terraform and infrastructure-as-code experience across multiple environments, with gitops and least-privilege practices.
- experience building ci/cd pipelines for infrastructure-as-code with automated plan/apply,
code review, policy-as-code, drift detection, and safe rollout.
- significant production experience with kubernetes, preferably gke, and helm-based workload deployment.
- strong cloud and l3/l4-l7 networking fundamentals, including vpcs, routing, load balancing, dns, tls, and interconnects.
- hands-on experience with prometheus, thanos, grafana, loki, tempo, and alertmanager for metrics, logs, traces, and alerting.
- operator-level familiarity with postgresql and message brokers such as rabbitmq or redpanda.
- understanding of sre practices, platform-as-a-product principles, incident management, capacity planning, and error budgets.
- willingness to participate in a follow-the-sun on-call rotation from apac hours and work effectively in a distributed, async-first environment.
- preferred experience includes opa/conftest, checkov, tflint, atlantis, terraform state and module management, backstage, tilt, alloy, rootly, go, linux, debian, ubuntu, docker, containerd, security and compliance, soc 2, secrets management, audit logging, and regulated fintech or low-latency systems.
benefits
- competitive salary and stock options.
- health benefits.
- one-time usd $500 new-hire home-office setup allowance.
- usd $150 monthly stipend through a brex card.
- globally distributed, async-first work environment with a follow-the-sun on-call model from apac hours.
📌 Senior devops engineer (México)
🏢 Alpaca
📍 México