Site Reliability Engineer (SRE) Regional Multi Project Platform (San Pedro Tlaquepaque)

Site Reliability Engineer (SRE) Regional Multi Project Platform (San Pedro Tlaquepaque)

27 ago
|
Qualcomm
|
San Pedro Tlaquepaque

27 ago

Qualcomm

San Pedro Tlaquepaque

## nCompany:nnQUALCOMM SEMICONDUCTORES Y SISTEMAS AVANZADOS DE BAJA CALIFORNIAnn## Job Area:nnEngineering Group, Engineering Group u003e Software EngineeringnnGeneral Summary:nnCloud Infrastructure u0026 Infrastructure as Codenn * Design, build, and manage cloud infrastructure with a primary focus on AWS, integrated with OpenStack environmentsn * Build and maintain Infrastructure as Code using:n * Terraformn * Ansiblen * Kubernetes (manifests / Helm)nnn * Design infrastructure solutions for:n * Scalabilityn * High availabilityn * Performancen * Reliabilityn * Cost efficiencynnn * Implement redundancy, failover, and disasteru2011recovery patterns across services and regionsn * Perform capacity planning based on performance metrics, usage trends, and utilization datannnnKubernetes u0026 Platform Reliabilitynn * Operate and scale production Kubernetes clusters in largeu2011scale environmentsn * Partner with development and QA teams to:n * Improve system reliability and resiliencyn * Automate scalability and availability mechanismsnnn * Apply SRE principles including:n * Service reliability ownershipn * Proactive failure preventionn * Continuous improvement of operational processesnnn * Support microservicesu2011based and distributed system architecturesnnnnCI/CD, Automation u0026 Operational Excellencenn * Manage and evolve CI/CD pipelines (e.g., Jenkins)n * Automate infrastructure provisioning, configuration, and lifecycle managementn * Write, maintain, and improve runbooks for operational processesn * Build automation to reduce manual intervention and operational toiln * Plan and execute infrastructure upgrades and maintenance activitiesn * Proactively identify and address technical and infrastructure debtnnnnData Platforms u0026 Streaming Systemsnn * Operate, tune, and scale data and streaming platforms, including:n * Kafka, Zookeepern * NiFin * Elasticsearchn * MySQL, Verticannn * Diagnose and resolve performance and stability issues across data pipelinesn * Ensure data platform reliability, throughput, and resilience at scalennnnAIu2011Assisted SRE u0026 Intelligent Automationnn * Design and maintain knowledgeu2011driven automated runbooks and operational botsn * Develop AIu2011assisted operational workflows, including:n * Incident analysis and summarizationn * Intelligent diagnostics and remediation suggestionsn * Automation of repetitive operational decisionu2011makingnnn * Work with LLMu2011based agent frameworks (e.g.,



Claude Agent SDK or similar):n * Integrate agents with logs, metrics, monitoring, and internal toolsn * Implement guardu2011railed, controlledu2011action automation for production usennn * Research and propose new concepts, tools, and AIu2011driven approaches to improve reliability and efficiencynnnnMonitoring, Reliability u0026 Incident Managementnn * Design and operate monitoring and observability systems using:n * Prometheusn * Grafanan * ELK stacknnn * Improve alert quality, signalu2011tou2011noise ratio, and troubleshooting efficiencyn * Lead incident response activities, root cause analysis, and postu2011incident reviewsn * Support software engineers in debugging complex production issues across distributed systemsn * Embed reliability, automation, and operational readiness into system designnnnnExperience Requirednn * Extensive experience operating largeu2011scale distributed cloud systemsn * Handsu2011on experience with AWS in production environmentsn * Direct experience working with OpenStackn * Strong Linux background in largeu2011scale SaaS or production systemsn * Ability to:n * Maintain and improve existing missionu2011critical systemsn * Prioritize and systematically reduce technical and infrastructure debtnnn * Strong understanding of designing for operational excellence, not just greenfield solutionsnnnnRequired Skillsnn * Programming: Strong experience with Python and/or Gon * Cloud u0026 IaC: Terraform, Ansible, CloudFormation or equivalentn * Containers: Kubernetes (production experience)n * CI/CD: Jenkins and modern CI/CD practicesn * Data u0026 Streaming: Kafka, NiFi, Elasticsearch, MySQL, Vertica, Zookeepern * Observability: Prometheus, Grafana, ELKn * Infrastructure: Nginx,



Linux internalsn * AI / Automation (advantage):n * Experience integrating AI or LLMs into operational workflowsn * Familiarity with agentu2011based automation conceptsnnnnExperience Guidelinesnn3+ years in:nn * overall experience managing infrastructuren * Linux administration in largeu2011scale environmentsn * operating production systems on AWS and/or OpenStackn * managing Kubernetes in productionn * using infrastructure as coden * working with CI/CD systemsnnnnMinimum Qualifications:nnu2022 Bacheloru0027s degree in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience. nOR nMasteru0027s degree in Engineering, Information Systems, Computer Science, or related field and 1+ year of Software Engineering or related work experience. nOR nPhD in Engineering, Information Systems, Computer Science, or related field. n nu2022 2+ years of academic or work experience with Programming Language such as C, C++, Java, Python, etc.nnApplicants: Qualcomm is an equal opportunity employer. If you are an individual with a disability and need an accommodation during the application/hiring process, rest assured that Qualcomm is committed to providing an accessible process. You may e-mail or call Qualcommu0027s toll-free number found here. Upon request, Qualcomm will provide reasonable accommodations to support individuals with disabilities to be able participate in the hiring process. Qualcomm is also committed to making our workplace accessible for individuals with disabilities. (Keep in mind that this email address is used to provide reasonable accommodations for individuals with disabilities. We will not respond here to requests for updates on applications or resume inquiries).nnQualcomm expects its employees to abide by all applicable policies and procedures, including but not limited to security and other requirements regarding protection of Company confidential information and other confidential and/or proprietary information, to the extent those requirements are permissible under applicable law.nnTo all Staffing and Recruiting Agencies: Our Careers Site is only for individuals seeking a job at Qualcomm. Staffing and recruiting agencies and individuals being represented by an agency are not authorized to use this site or to submit profiles, applications or resumes, and any such submissions will be considered unsolicited. Qualcomm does not accept unsolicited resumes or applications from agencies. Please do not forward resumes to our jobs alias, Qualcomm employees or any other company location. Qualcomm is not responsible for any fees related to unsolicited resumes/applications.nnIf you would like more information about this role, please contact Qualcomm Careers.n

📌 Site Reliability Engineer (SRE) Regional Multi Project Platform (San Pedro Tlaquepaque)
🏢 Qualcomm
📍 San Pedro Tlaquepaque

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (sre) regional multi project platform (san pedro tlaquepaque) / san pedro tlaquepaque

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (sre) regional multi project platform (san pedro tlaquepaque) / san pedro tlaquepaque