Site reliability engineer (sre) regional multi project platform (Zapopan)

Site reliability engineer (sre) regional multi project platform (Zapopan)

28 ago
|
Qualcomm
|
Zapopan

28 ago

Qualcomm

Zapopan

##
Company: QUALCOMM SEMICONDUCTORES Y SISTEMAS AVANZADOS DE BAJA CALIFORNIA ## Job Area: Engineering Group, Engineering Group u003e Software Engineering General Summary: Cloud Infrastructure u0026 Infrastructure as Code
* Design, build, and manage cloud infrastructure with a primary focus on AWS, integrated with Open Stack environments
* Build and maintain Infrastructure as Code using:
* Terraform
* Ansible
* Kubernetes (manifests / Helm)
* Design infrastructure solutions for:
* Scalability
* High availability
* Performance
* Reliability
* Cost efficiency
* Implement redundancy, failover, and disasteru2011recovery patterns across services and regions
* Perform capacity planning based on performance metrics, usage trends, and utilization data Kubernetes u0026 Platform Reliability
* Operate and scale production Kubernetes clusters in largeu2011scale environments
* Partner with development and QA teams to:
* Improve system reliability and resiliency
* Automate scalability and availability mechanisms
* Apply SRE principles including:
* Service reliability ownership
* Proactive failure prevention
* Continuous improvement of operational processes
* Support microservicesu2011based and distributed system architectures CI/CD, Automation u0026 Operational Excellence
* Manage and evolve CI/CD pipelines (e.g., Jenkins)
* Automate infrastructure provisioning, configuration, and lifecycle management
* Write, maintain, and improve runbooks for operational processes
* Build automation to reduce manual intervention and operational toil
* Plan and execute infrastructure upgrades and maintenance activities
* Proactively identify and address technical and infrastructure debt Data Platforms u0026 Streaming Systems
* Operate, tune, and scale data and streaming platforms, including:
* Kafka, Zookeeper
* Ni Fi
* Elasticsearch
* My SQL, Vertica
* Diagnose and resolve performance and stability issues across data pipelines




* Ensure data platform reliability, throughput, and resilience at scale AIu2011 Assisted SRE u0026 Intelligent Automation
* Design and maintain knowledgeu2011driven automated runbooks and operational bots
* Develop AIu2011assisted operational workflows, including:
* Incident analysis and summarization
* Intelligent diagnostics and remediation suggestions
* Automation of repetitive operational decisionu2011making
* Work with LLMu2011based agent frameworks (e.g., Claude Agent SDK or similar):
* Integrate agents with logs, metrics, monitoring, and internal tools
* Implement guardu2011railed, controlledu2011action automation for production use
* Research and propose new concepts, tools, and AIu2011driven approaches to improve reliability and efficiency Monitoring, Reliability u0026 Incident Management
* Design and operate monitoring and observability systems using:
* Prometheus
* Grafana
* ELK stack
* Improve alert quality, signalu2011tou2011noise ratio, and troubleshooting efficiency
* Lead incident response activities, root cause analysis, and postu2011incident reviews
* Support software engineers in debugging complex production issues across distributed systems
* Embed reliability, automation, and operational readiness into system design Experience Required
* Extensive experience operating largeu2011scale distributed cloud systems
* Handsu2011on experience with AWS in production environments
* Direct experience working with Open Stack
* Strong Linux background in largeu2011scale Saa S or production systems
* Ability to:




* Maintain and improve existing missionu2011critical systems
* Prioritize and systematically reduce technical and infrastructure debt
* Strong understanding of designing for operational excellence, not just greenfield solutions Required Skills
* Programming: Strong experience with Python and/or Go
* Cloud u0026 Ia C: Terraform, Ansible, Cloud Formation or equivalent
* Containers: Kubernetes (production experience)
* CI/CD: Jenkins and modern CI/CD practices
* Data u0026 Streaming: Kafka, Ni Fi, Elasticsearch, My SQL, Vertica, Zookeeper
* Observability: Prometheus, Grafana, ELK
* Infrastructure: Nginx, Linux internals
* AI / Automation (advantage):
* Experience integrating AI or LLMs into operational workflows
* Familiarity with agentu2011based automation concepts Experience Guidelines 3+ years in:
* overall experience managing infrastructure
* Linux administration in largeu2011scale environments
* operating production systems on AWS and/or Open Stack
* managing Kubernetes in production
* using infrastructure as code
* working with CI/CD systems Minimum Qualifications: u2022 Bacheloru0027s degree in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience.
OR
Masteru0027s degree in Engineering, Information Systems, Computer Science, or related field and 1+ year of Software Engineering or related work experience.
OR
Ph D in Engineering, Information Systems, Computer Science, or related field.

u2022 2+ years of academic or work experience with Programming Language such as C, C++, Java, Python, etc. Applicants: Qualcomm is an equal opportunity employer. If you are an individual with a disability and need an accommodation during the application/hiring process, rest assured that Qualcomm is committed to providing an accessible process. You may e-mail

📌 Site reliability engineer (sre) regional multi project platform (Zapopan)
🏢 Qualcomm
📍 Zapopan

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (sre) regional multi project platform (zapopan) / zapopan

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (sre) regional multi project platform (zapopan) / zapopan