07 ago
|
Tata Consultancy Services
|
Nezahualcóyotl
07 ago
Tata Consultancy Services
Nezahualcóyotl
Tata Consultancy Services is an equal opportunity employer, our commitment to diversity & inclusion drives our efforts to provide equal opportunity to all candidates who meet our required knowledge & competency needs, irrespective of any socio-economic background, race, color, national origin, religion, sex, gender identity/expression, age, marital status, disability, sexual orientation or any others. We encourage anyone interested to build a career in TCS to participate in our recruitment & selection process. TCS is seeking skilled professionals to join our team as SRE. Technical/Functional Skills: Observability & Monitoring Develop proactive alerting and dashboarding strategies to detect and resolve issues before they impact customers and store operations Define and manage Service Level Objectives (SLOs), Service Level Agreements (SLAs), and error budgets for critical store applications Lead critical incident recovery and postmortem processes to drive continuous improvement Performance & Reliability Engineering Identify and eliminate bottlenecks in development and deployment workflows to improve lead time and reduce change failure rates Partner with development teams to embed Site Reliability Engineering (SRE)
principles into the software development lifecycle Support and optimize applications deployed across CVS retail and pharmacy locations Collaborate with infrastructure and store operations teams to ensure high availability and performance of store systems Microservices & Deployments Champion containerization and orchestration using Open Shift and Kubernetes in hybrid cloud environments Leverage CI/CD pipelines to enable automated deployments at scale Understanding of microservices architecture Minimum Qualifications: 5+ years of experience in SRE, Dev Ops, or related technology roles 3+ years of experience in delivering software in a large-scale environment with reliability and resilience concepts (multi-region, multi-cloud, containerization, etc.) 2+ years of experience with programming languages/frameworks 2+ years of experience on Cloud Technologies (AWS, Microsoft Azure, Google Cloud), Microservices concepts, and capabilities like Rancher, Docker, Kubernetes, and web API's 2+ years of experience with source control and continuous integration tools like Git Hub, Bitbucket, or Jenkins Experience with observability and monitoring tools such as Splunk, Dynatrace, Datadog, Prometheus, Grafana, etc. Proficiency in scripting and automation frameworks Understanding of microservices architecture and cloud-native technologies Experience in Incident Management, Change Management, Infrastructure Support, and Problem Management concepts and processes Excellent interpersonal and communication skills, including the ability to engage technical and non-technical stakeholder**Work modality: Hybrid** Candidate must be located in or willing to relocate to Querétaro, CDMX, Monterrey, Guadalajara, it will be requested to attend office at least 3 days per week. Boost your career and send your resume to: alejandra.galicia@
📌 Site reliability engineer (Nezahualcóyotl)
🏢 Tata Consultancy Services
📍 Nezahualcóyotl