29 ago
|
Qualcomm
|
Saltillo
nCompany:nnQUALCOMM SEMICONDUCTORES Y SISTEMAS AVANZADOS DE BAJA CALIFORNIAnn## Job Area:nnEngineering Group, Engineering Group u003e Software EngineeringnnGeneral Summary:nnCloud Infrastructure u0026 Infrastructure as Codenn Design, build, and manage cloud infrastructure with a primary focus on AWS, integrated with OpenStack environmentsn Build and maintain Infrastructure as Code using:n Terraformn Ansiblen Kubernetes (manifests / Helm)nnn Design infrastructure solutions for:n Scalabilityn High availabilityn Performancen Reliabilityn Cost efficiencynnn Implement redundancy, failover, and disasteru2011recovery patterns across services and regionsn Perform capacity planning based on performance metrics, usage trends, and utilization datannnnKubernetes u0026 Platform Reliabilitynn Operate and scale production Kubernetes clusters in largeu2011scale environmentsn Partner with development and QA teams to:n Improve system reliability and resiliencyn Automate scalability and availability mechanismsnnn Apply SRE principles including:n Service reliability ownershipn Proactive failure preventionn Continuous improvement of operational processesnnn Support microservicesu2011based and distributed system architecturesnnnnCI/CD, Automation u0026 Operational Excellencenn Manage and evolve CI/CD pipelines (e.g., Jenkins)n Automate infrastructure provisioning, configuration, and lifecycle managementn Write, maintain, and improve runbooks for operational processesn Build automation to reduce manual intervention and operational toiln Plan and execute infrastructure upgrades and maintenance activitiesn Proactively identify and address technical and infrastructure debtnnnnData Platforms u0026 Streaming Systemsnn Operate, tune,
and scale data and streaming platforms, including:n Kafka, Zookeepern NiFin Elasticsearchn MySQL, Verticannn Diagnose and resolve performance and stability issues across data pipelinesn Ensure data platform reliability, throughput, and resilience at scalennnnAIu2011Assisted SRE u0026 Intelligent Automationnn Design and maintain knowledgeu2011driven automated runbooks and operational botsn Develop AIu2011assisted operational workflows, including:n Prometheusn Grafanan ELK stacknnn Improve alert quality, signalu2011tou2011noise ratio, and troubleshooting efficiencyn Lead incident response activities, root cause analysis, and postu2011incident reviewsn Support software engineers in debugging complex production issues across distributed systemsn Embed reliability, automation, and operational readiness into system designnnnnExperience Requirednn Extensive experience operating largeu2011scale distributed cloud systemsn Handsu2011on experience with AWS in production environmentsn Direct experience working with OpenStackn Strong Linux background in largeu2011scale SaaS or production systemsn Ability to:n Maintain and improve existing missionu2011critical systemsn Prioritize and systematically reduce technical and infrastructure debtnnn Strong understanding of designing for operational excellence, not just greenfield solutionsnnnnRequired Skillsnn Programming:
Strong experience with Python and/or Gon Cloud u0026 IaC: Terraform, Ansible, CloudFormation or equivalentn Containers: Kubernetes (production experience)n CI/CD: Jenkins and modern CI/CD practicesn Data u0026 Streaming: Kafka, NiFi, Elasticsearch, MySQL, Vertica, Zookeepern Observability: Prometheus, Grafana, ELKn Infrastructure: Nginx, Linux internalsn AI / Automation (advantage):n Experience integrating AI or LLMs into operational workflowsn Familiarity with agentu2011based automation conceptsnnnnExperience Guidelinesnn3+ years in:nn overall experience managing infrastructuren Linux administration in largeu2011scale environmentsn operating production systems on AWS and/or OpenStackn managing Kubernetes in productionn using infrastructure as coden working with CI/CD systemsnnnnMinimum Qualifications:nnu2022 Bacheloru0027s degree in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience. nOR nMasteru0027s degree in Engineering, Information Systems, Computer Science, or related field and 1+ year of Software Engineering or related work experience. nOR nPhD in Engineering, Information Systems, Computer Science, or related field. n nu2022 2+ years of academic or work experience with Programming Language such as C, C++, Java, Python, etc.nnQualcomm expects its employees to abide by all applicable policies and procedures, including but not limited to security and other requirements regarding protection of Company confidential information and other confidential and/or proprietary information, to the extent those requirements are permissible under applicable law.
📌 Site Reliability Engineer (SRE) Regional Multi Project Platform (Saltillo)
🏢 Qualcomm
📍 Saltillo