22 ago
|
Epam Systems
|
México
22 ago
Epam Systems
México
We are looking for an experienced **Site Reliability Engineer (SRE)** to take a leadership role in ensuring the stability, scalability, and performance of our cloud infrastructure on **Google Cloud Platform (GCP)**. As an SRE, you will be at the forefront of optimizing system reliability, automating processes, and collaborating with engineering teams to enhance operational excellence. If you're passionate about **infrastructure-as-code, automation, and building resilient systems**, we’d love to hear from you.
**Responsibilities**
- Lead reliability initiatives to optimize system performance, scalability, and cost efficiency
- Manage and participate in on-call rotations, providing 24/7 support for critical infrastructure
- Troubleshoot incidents, conduct root cause analysis (RCA), and implement long-term solutions
- Deploy and manage microservices in alignment with release cycles
- Design and maintain infrastructure-as-code solutions using Terraform
- Collaborate with development teams to improve system reliability, performance, and cloud resource management
- Oversee incident response and ticket management using ServiceNow and Jira
- Maintain and expand internal knowledge bases on infrastructure and monitoring
**Requirements**:
- 5+ years of experience in SRE, DevOps, or system administration roles
- Expertise in Google Cloud Platform (GCP) and cloud-native architectures
- Hands-on experience with incident management and monitoring tools (ServiceNow, Cloud Monitoring, etc.)
- Strong debugging and problem-solving skills for complex technical issues
- Proficiency in GitHub and infrastructure-as-code best practices
- Strong communication and teamwork skills, with a proactive mindset
**Nice to have**
- Experience with Kubernetes and containerization technologies
- Deep understanding of CI/CD pipelines and related tools
- Familiarity with Prometheus, Grafana, Catchpoint, and ELK for monitoring and logging
**We offer**
- Career plan and real growth opportunities
- Unlimited access to LinkedIn learning solutions
- International Mobility Plan within 25 countries
- Constant training, mentoring, online corporate courses, eLearning and more
- English classes with a certified teacher
- Support for employee’s initiatives (Algorithms club, toastmasters, agile club and more)
- Enjoyable working environment (Gaming room, napping area, amenities, events, sport teams and more)
- Flexible work schedule and dress code
- Collaborate in a multicultural environment and share best practices from around the globe
- Hired directly by EPAM & 100% under payroll
- Law benefits (IMSS, INFONAVIT, 25% vacation bonus)
- Major medical expenses insurance: Life, Major medical expenses with dental & visual coverage (for the employee and direct family members)
- 13 % employee savings fund, capped to the law limit
- Grocery coupons
- 30 days December bonus
- Employee Stock Purchase Plan
- 12 vacations days plus 4 floating days
- Official Mexican holidays, plus 5 extra holidays (Maundry Thursday and Friday, November 2nd, December 24th & 31st)
- Monthly non-taxable amount for the electricity and internet bills
EPAM is a leading general provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.
📌 Lead Site Reliability Engineer (México)
🏢 Epam Systems
📍 México