09 ago
|
Qualcomm
|
Tijuana
Company
QUALCOMM SEMICONDUCTORES Y SISTEMAS AVANZADOS DE BAJA CALIFORNIA Software EngineeringGeneral Summary
Lead design, build, and manage cloud infrastructure and operate large‐scale distributed systems at Qualcomm.
Work across cloud, Kubernetes, CI/CD, data platforms, AI‐assisted SRE, and monitoring to deliver reliable, scalable, cost‐efficient services.Cloud Infrastructure & Infrastructure as CodeDesign, build, and manage cloud infrastructure with a primary focus on AWS, integrated with Open Stack environmentsBuild and maintain Infrastructure as Code using Terraform, Ansible, and Kubernetes manifests/HelmDesign infrastructure solutions for scalability, high availability, performance, reliability, and cost efficiencyImplement redundancy, failover, and disaster‐recovery patterns across services and regionsPerform capacity planning based on performance metrics, usage trends, and utilization dataKubernetes & Platform ReliabilityOperate and scale production Kubernetes clusters in large‐scale environmentsPartner with development and QA teams to improve system reliability, resiliency, and automate scalability and availability mechanismsApply SRE principles, including service reliability ownership, proactive failure prevention, continuous improvement of operational processes, and support microservices‐based and distributed system architecturesCI/CD, Automation & Operational ExcellenceManage and evolve CI/CD pipelines (e.g., Jenkins)Automate infrastructure provisioning, configuration, and lifecycle managementWrite, maintain, and improve runbooks for operational processesBuild automation to reduce manual intervention and operational toilPlan and execute infrastructure upgrades and maintenance activitiesProactively identify and address technical and infrastructure debtData Platforms & Streaming SystemsOperate, tune, and scale data and streaming platforms including Kafka, Zookeeper, Ni Fi,
Elasticsearch, My SQL, and VerticaDiagnose and resolve performance and stability issues across data pipelinesEnsure data platform reliability, throughput, and resilience at scaleAI‐Assisted SRE & Intelligent AutomationDesign and maintain knowledge‐driven automated runbooks and operational botsDevelop AI‐assisted operational workflows, including incident analysis and summarization, intelligent diagnostics and remediation suggestions, and automation of repetitive operational decision‐makingWork with LLM‐based agent frameworks (e.g., Claude Agent SDK), integrating agents with logs, metrics, monitoring, and internal tools, implementing guard‐railed, controlled‐action automation for production use, and researching new AI‐driven approaches to improve reliability and efficiencyMonitoring, Reliability & Incident ManagementDesign and operate monitoring and observability systems using Prometheus, Grafana, and ELK stackImprove alert quality, signal‐to‐noise ratio, and troubleshooting efficiencyLead incident response activities, root cause analysis, and post‐incident reviewsSupport software engineers in debugging complex production issues across distributed systemsEmbed reliability, automation,
and operational readiness into system designExperience RequiredExtensive experience operating large‐scale distributed cloud systemsHands‐on experience with AWS in production environmentsDirect experience working with Open StackStrong Linux background in large‐scale Saa S or production systemsAbility to maintain and improve existing mission‐critical systems, prioritize and systematically reduce technical and infrastructure debt, and design for operational excellenceRequired SkillsProgramming: Strong experience with Python and/or GoCloud & Ia C: Terraform, Ansible, Cloud Formation or equivalentContainers: Kubernetes (production experience)CI/CD: Jenkins and modern CI/CD practicesData & Streaming: Kafka, Ni Fi, Elasticsearch, My SQL, Vertica, ZookeeperObservability: Prometheus, Grafana, ELKInfrastructure: Nginx, Linux internalsAI / Automation (advantage): Experience integrating AI or LLMs into operational workflows, familiarity with agent‐based automation conceptsExperience Guidelines3+ years overall experience managing infrastructure3+ years Linux administration in large‐scale environments3+ years operating production systems on AWS and/or Open Stack3+ years managing Kubernetes in production3+ years using infrastructure as code3+ years working with CI/CD systemsMinimum QualificationsBachelor's degree in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience.or Master's degree in Engineering, Information Systems, Computer Science, or related field and 1+ year of Software Engineering or related work experience.or Ph D in Engineering, Information Systems, Computer Science, or related field.2+ years of academic or work experience with programming languages such as C, C++, Java, Python, etc.Applicants
Qualcomm is an equal opportunity employer.
If you are an individual with a disability and need an accommodation during the application/hiring process, Qualcomm is committed to providing an accessible process.
For accommodations, contact disability‐accomodations@ or call Qualcomm's toll‐free number.
📌 Site Reliability Engineer (Sre) – Regional Multi Project Platform (Tijuana)
🏢 Qualcomm
📍 Tijuana