Site Reliability Engineer (Sre) – Regional Multi Project Platform (Tijuana)

Site Reliability Engineer (Sre) – Regional Multi Project Platform (Tijuana)

01 ago
|
Qualcomm
|
Tijuana

01 ago

Qualcomm

Tijuana

QUALCOMM SEMICONDUCTORES Y SISTEMAS AVANZADOS DE BAJA CALIFORNIAJob AreaEngineering Group, Engineering Group > Software EngineeringGeneral SummaryCloud Infrastructure & Infrastructure as CodeDesign, build, and manage cloud infrastructure with a primary focus on AWS, integrated with OpenStack environmentsBuild and maintain Infrastructure as Code using:TerraformAnsibleKubernetes (manifests / Helm)Design infrastructure solutions for:ScalabilityHigh availabilityPerformanceReliabilityCost efficiencyImplement redundancy, failover, and disaster‐recovery patterns across services and regionsPerform capacity planning based on performance metrics, usage trends, and utilization dataKubernetes & Platform ReliabilityOperate and scale production Kubernetes clusters in large‐scale environmentsPartner with development and QA teams to:Improve system reliability and resiliencyAutomate scalability and availability mechanismsApply SRE principles including:Service reliability ownershipProactive failure preventionContinuous improvement of operational processesSupport microservices‐based and distributed system architecturesCI/CD, Automation & Operational ExcellenceManage and evolve CI/CD pipelines (e.g., Jenkins)Automate infrastructure provisioning, configuration, and lifecycle managementWrite, maintain, and improve runbooks for operational processesBuild automation to reduce manual intervention and operational toilPlan and execute infrastructure upgrades and maintenance activitiesProactively identify and address technical and infrastructure debtData Platforms & Streaming SystemsOperate, tune, and scale data and streaming platforms, including:Kafka, ZookeeperNiFiElasticsearchMySQL, VerticaDiagnose and resolve performance and stability issues across data pipelinesEnsure data platform reliability, throughput,



and resilience at scaleAI‐Assisted SRE & Intelligent AutomationDesign and maintain knowledge‐driven automated runbooks and operational botsDevelop AI‐assisted operational workflows, including:Incident analysis and summarizationIntelligent diagnostics and remediation suggestionsAutomation of repetitive operational decision‐makingWork with LLM‐based agent frameworks (e.g., Claude Agent SDK or similar):Integrate agents with logs, metrics, monitoring, and internal toolsImplement guard‐railed, controlled‐action automation for production useResearch and propose new concepts, tools, and AI‐driven approaches to improve reliability and efficiencyMonitoring, Reliability & Incident ManagementDesign and operate monitoring and observability systems using:PrometheusGrafanaELK stackImprove alert quality, signal‐to‐noise ratio, and troubleshooting efficiencyLead incident response activities, root cause analysis, and post‐incident reviewsSupport software engineers in debugging complex production issues across distributed systemsEmbed reliability, automation, and operational readiness into system designExperience RequiredExtensive experience operating large‐scale distributed cloud systemsHands‐on experience with AWS in production environmentsDirect experience working with OpenStackStrong Linux background in large‐scale SaaS or production systemsAbility to:Maintain and improve existing mission‐critical systemsPrioritize and systematically reduce technical and infrastructure debtStrong understanding of designing for operational excellence, not just greenfield solutionsRequired SkillsProgramming:



Strong experience with Python and/or GoCloud & IaC: Terraform, Ansible, CloudFormation or equivalentContainers: Kubernetes (production experience)CI/CD: Jenkins and modern CI/CD practicesData & Streaming: Kafka, NiFi, Elasticsearch, MySQL, Vertica, ZookeeperObservability: Prometheus, Grafana, ELKInfrastructure: Nginx, Linux internalsAI / Automation (advantage):Experience integrating AI or LLMs into operational workflowsFamiliarity with agent‐based automation conceptsExperience Guidelines3+ years in:overall experience managing infrastructureLinux administration in large‐scale environmentsoperating production systems on AWS and/or OpenStackmanaging Kubernetes in productionusing infrastructure as codeworking with CI/CD systemsMinimum QualificationsBachelor's degree in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience.Master's degree in Engineering, Information Systems, Computer Science, or related field and 1+ year of Software Engineering or related work experience.PhD in Engineering, Information Systems, Computer Science, or related field.2+ years of academic or work experience with Programming Language such as C, C++, Java, Python, etc.Equal Opportunity Employer StatementApplicants: Qualcomm is an equal opportunity employer.
If you are an individual with a disability and need an accommodation during the application/hiring process, rest assured that Qualcomm is committed to providing an accessible process.Qualcomm will provide reasonable accommodations to support individuals with disabilities to be able participate in the hiring process.
You may e-mail or call Qualcomm's toll-free number found here.Qualcomm is also committed to making our workplace accessible for individuals with disabilities.
#J-*****-Ljbffr

📌 Site Reliability Engineer (Sre) – Regional Multi Project Platform (Tijuana)
🏢 Qualcomm
📍 Tijuana

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (sre) – regional multi project platform (tijuana) / tijuana

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (sre) – regional multi project platform (tijuana) / tijuana