07 ago
|
Petco
|
Santiago de Querétaro
07 ago
Petco
Santiago de Querétaro
Responsibilities
- Serve as a technical owner for production support, ensuring the stability, availability, and performance of Python backend services, React Native mobile applications, GraphQL APIs, PostgreSQL databases, and AWS cloud infrastructure.
- Lead the investigation and resolution of production incidents, service disruptions, performance degradation, and customer-reported issues.
- Analyze application logs, system metrics, distributed traces, and database activity to quickly identify root causes and restore service.
- Troubleshoot complex issues across application, infrastructure, database, network, and integration layers.
- Perform detailed log analysis using Datadog, CloudWatch, and other observability tools to identify failure patterns, bottlenecks, and system anomalies.
- Investigate data-related issues by executing SQL queries, analyzing database performance, validating data integrity, and troubleshooting PostgreSQL and DynamoDB-related problems.
- Develop and deploy bug fixes, hotfixes, and permanent corrective actions to prevent incident recurrence.
- Participate in on-call rotations and act as a senior escalation point for critical production incidents.
- Coordinate incident response activities, provide timely stakeholder communication, and drive service restoration efforts.
- Conduct root cause analysis (RCA) for major incidents and document findings, corrective actions, and preventive measures.
- Create and maintain operational runbooks, troubleshooting guides, support documentation, and knowledge base articles.
- Improve monitoring, alerting, logging, and observability capabilities to proactively detect issues before they impact customers.
- Partner with development teams to prioritize production defects, technical debt remediation, and system reliability improvements.
- Support application releases, production deployments, and change management activities,
ensuring successful implementation and post-deployment validation.
- Identify opportunities for automation to reduce manual support effort and improve operational efficiency.
- Mentor engineers on production support best practices, troubleshooting techniques, and operational excellence.
Education / Experience
- Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience.
- 5+ years of experience supporting and troubleshooting large-scale production systems.
- Strong hands-on experience with Python, GraphQL, PostgreSQL, React Native, and AWS cloud services.
- Proven experience diagnosing and resolving production incidents in distributed cloud-native environments.
- Strong experience analyzing application logs, metrics, traces, and monitoring data using Datadog, CloudWatch, or similar observability platforms.
- Advanced SQL skills with the ability to investigate data issues, analyze query performance, and troubleshoot database-related incidents.
- Experience supporting AWS services such as Lambda, ECS, DynamoDB, API Gateway, S3, and CloudWatch.
- Strong understanding of incident management, escalation procedures, problem management, and root cause analysis processes.
- Experience developing production bug fixes, hotfixes, and permanent corrective actions in a fast-paced environment.
- Utilize AI-assisted development tools to improve troubleshooting efficiency, generate code fixes, automate repetitive support tasks, and reduce mean time to resolution (MTTR).
- Excellent analytical, troubleshooting, and problem-solving skills with the ability to quickly isolate and resolve complex issues.
- Strong verbal and written communication skills with the ability to communicate technical issues clearly to both technical and business stakeholders.
- Experience working in a 24x7 production support environment and participating in on-call rotations.
#J-18808-Ljbffr
📌 Sr. Software Engineer (Santiago de Querétaro)
🏢 Petco
📍 Santiago de Querétaro