04 sep
|
Serve Robotics
|
Venustiano Carranza
04 sep
Serve Robotics
Venustiano Carranza
At Serve Robotics, we're reimagining how things move in cities. It's designed to take deliveries away from congested streets, make deliveries available to more people, and benefit local businesses. The Serve fleet has been delighting merchants, customers, and pedestrians along the way in Los Angeles, Miami, Dallas, Atlanta and Chicago while doing commercial deliveries.
We are tech industry veterans in software, hardware, and design who are pooling our skills to build the future we want to live in. We are solving real-world problems leveraging robotics, machine learning and computer vision, among other disciplines, with a mindful eye towards the end-to-end user experience. Our team is agile, diverse, and driven.
The Senior Reliability Operations
Engineer leads operational reliability by region owning incident response, escalations, and Tier 2 support for robotic and cloud systems. This role drives the creation and improvement of runbooks, automations, and operational processes while coordinating closely with product engineering and SREs. This position serves as the regional incident lead, ensuring issues are resolved efficiently and communicated clearly to all stakeholders.
Responsibilities Serve as the primary incident lead during your region's daytime hours, coordinating technical investigations, centralizing communication, and engaging the appropriate engineering and SRE teams when escalation is required. Use metrics, logs, and tracing tools (Grafana/Prometheus, GCP Monitoring, OpenTelemetry) to proactively identify problems, validate system behavior, and support continuous improvement of detection mechanisms. Act as the central point of communication during active incidents, ensuring timely updates and clear routing to the correct product engineering and SRE stakeholders.
Participate in a shared weekend on-call rotation to help maintain operational coverage for production systems, responding to incidents and escalations as needed and coordinating with engineering teams when issues arise.
Qualifications Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent practical experience. 5+ years of professional experience in Reliability Operations, Site Reliability Engineering, DevOps, IT Operations, or a related technical support function. Demonstrated experience owning or participating in Tier 2 or Tier 3 technical investigations, including triage, log analysis, and structured escalation.
Experience supporting distributed systems, cloud-hosted services, or production operational environments. Strong proficiency with Linux, including navigating systems, reviewing logs, and performing diagnostics.
Experience writing, executing, and maintaining runbooks, automations, and operational workflows. Ability to interpret metrics, logs, and traces using tools such as Grafana/Prometheus, Google Cloud Monitoring, and OpenTelemetry. Familiarity with modern cloud environments, preferably Google Cloud Platform (GCP), including basic debugging, permissions, and service-level triage.
Proficiency with Jira or similar platforms for ticketing and structured incident tracking. Strong collaboration skills when coordinating with product engineering, SRE, and general support teams. Familiarity with incident management tools such as PagerDuty, OpsGenie, Jira Service Management, or Grafana IRM.
Strong networking fundamentals, including experience diagnosing connectivity issues across distributed systems. Familiarity with Tailscale or similar zero-trust networking tools is a major plus.
Additional Information
As part of maintaining continuous operational coverage, this role also participates in a rotating weekend on-call schedule shared across the Reliability Operations team.
📌 Senior Production Engineering Specialist (Venustiano Carranza)
🏢 Serve Robotics
📍 Venustiano Carranza