30 ago
|
Qualcomm
|
Nuevo León
30 ago
Qualcomm
Nuevo León
##
Company:
QUALCOMM SEMICONDUCTORES Y SISTEMAS AVANZADOS DE BAJA CALIFORNIA
## Job Area:
Engineering Group, Engineering Group u003e Software Engineering
General Summary:
Cloud Infrastructure u**** Infrastructure as Code
* Design, build, and manage cloud infrastructure with a primary focus on AWS, integrated with Open Stack environments
* Build and maintain Infrastructure as Code using:
* Terraform
* Ansible
* Kubernetes (manifests / Helm)
* Design infrastructure solutions for:
* Scalability
* High availability
* Performance
* Reliability
* Cost efficiency
* Implement redundancy, failover, and disasteru****recovery patterns across services and regions
* Perform capacity planning based on performance metrics, usage trends, and utilization data
Kubernetes u**** Platform Reliability
* Operate and scale production Kubernetes clusters in largeu****scale environments
* Partner with development and QA teams to:
* Improve system reliability and resiliency
* Automate scalability and availability mechanisms
* Apply SRE principles including:
* Service reliability ownership
* Proactive failure prevention
* Continuous improvement of operational processes
* Support microservicesu****based and distributed system architectures
CI/CD, Automation u**** Operational Excellence
* Manage and evolve CI/CD pipelines (e.g., Jenkins)
* Automate infrastructure provisioning, configuration, and lifecycle management
* Write, maintain, and improve runbooks for operational processes
* Build automation to reduce manual intervention and operational toil
* Plan and execute infrastructure upgrades and maintenance activities
* Proactively identify and address technical and infrastructure debt
Data Platforms u**** Streaming Systems
* Operate, tune, and scale data and streaming platforms, including:
* Kafka, Zookeeper
* Ni Fi
* Elasticsearch
* My SQL, Vertica
* Diagnose and resolve performance and stability issues across data pipelines
* Ensure data platform reliability, throughput, and resilience at scale
AIu**** Assisted SRE u**** Intelligent Automation
* Design and maintain knowledgeu****driven automated runbooks and operational bots
* Develop AIu****assisted operational workflows, including:
* Incident analysis and summarization
* Intelligent diagnostics and remediation suggestions
* Automation of repetitive operational decisionu****making
* Work with LLMu****based agent frameworks (e.g., Claude Agent SDK or similar):
* Integrate agents with logs, metrics, monitoring, and internal tools
* Implement guardu****railed, controlledu****action automation for production use
* Research and propose new concepts, tools, and AIu****driven approaches to improve reliability and efficiency
Monitoring, Reliability u**** Incident Management
* Design and operate monitoring and observability systems using:
* Prometheus
* Grafana
* ELK stack
* Improve alert quality, signalu****tou****noise ratio, and troubleshooting efficiency
* Lead incident response activities, root cause analysis, and postu****incident reviews
* Support software engineers in debugging complex production issues across distributed systems
* Embed reliability, automation, and operational readiness into system design
Experience Required
* Extensive experience operating largeu****scale distributed cloud systems
* Handsu****on experience with AWS in production environments
* Direct experience working with Open Stack
* Strong Linux background in largeu****scale Saa S or production systems
* Ability to:
* Maintain and improve existing missionu****critical systems
* Prioritize and systematically reduce technical and infrastructure debt
* Strong understanding of designing for operational excellence, not just greenfield solutions
Required Skills
* Programming: Strong experience with Python and/or Go
* Cloud u**** Ia C: Terraform, Ansible, Cloud Formation or equivalent
* Containers: Kubernetes (production experience)
* CI/CD: Jenkins and modern CI/CD practices
* Data u**** Streaming: Kafka, Ni Fi, Elasticsearch, My SQL, Vertica, Zookeeper
* Observability: Prometheus, Grafana, ELK
* Infrastructure: Nginx, Linux internals
* AI / Automation (advantage):
* Experience integrating AI or LLMs into operational workflows
* Familiarity with agentu****based automation concepts
Experience Guidelines
3+ years in:
* overall experience managing infrastructure
* Linux administration in largeu****scale environments
* operating production systems on AWS and/or Open Stack
* managing Kubernetes in production
* using infrastructure as code
* working with CI/CD systems
Minimum Qualifications:
u**** Bacheloru****s degree in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience.
OR
Masteru****s degree in Engineering, Information Systems, Computer Science, or related field and 1+ year of Software Engineering or related work experience.
OR
Ph D in Engineering, Information Systems, Computer Science, or related field.
u******+ years of academic or work experience with Programming Language such as C, C++, Java, Python, etc.
Applicants: Qualcomm is an equal opportunity employer.
If you are an individual with a disability and need an accommodation during the application/hiring process, rest assured that Qualcomm is committed to providing an accessible process.
You may e-mail
📌 Site Reliability Engineer (Sre) Regional Multi Project Platform (Nuevo León)
🏢 Qualcomm
📍 Nuevo León