18 ago
|
LTIMindtree
|
Ciudad de México
18 ago
LTIMindtree
Ciudad de México
Job Title: Site Reliability Engineer (SRE) - Streaming Platform
Location Mexico, Brazil, CR
Work Mode Remote
Hiring Mode for Mexico, Brazil and CR FTE and as Contractor only Brazil
Salary: 90,000 MXN plus Benefits / 5,000 USD CR / 40USD Contractor
This is a critical role with a wide range of responsibilities, including:
Analyze and improve system design to reduce failure modes and promote self-healing systems.
Establish and maintain robust systems that facilitate observability, encompassing logging, monitoring, distributed tracing, alerting, and offline test tools.
Work with development partners to shape the architecture, design, and implementation of new and existing systems to enhance their reliability, performance, efficiency, and scalability.
Ability to work both independently and as part of a geographically dispersed, yet integrated, team.
Collaborate with service engineers to establish Service Level Agreements (SLAs) and Service Level Objectives (SLOs) for backend services.
Identify indications or cues that demonstrate the effectiveness of an application and possess the knowledge to improve or repair its performance.
Assess options and suggest solutions when information is limited or unclear. This position requires a level of comfort and confidence in dealing with uncertain situations.
Work seamlessly within a team while effectively managing individual tasks.
Respond to emerging incidents, solve critical issues,
and follow through with a plan for resolution or future mitigation.
Act as an SME on the Engineering Operations team, partnering with backend services teams and application teams to overcome challenges across all platforms where we stream our service.
Qualities & Experience We're Seeking
We believe the right individual will have the following skills and experience to be successful in the role:
5+ years of experience in software development.
Degree in Computer Science or a related field, or equivalent work experience.
Solid engineering and coding skills, strong data structure knowledge, and the ability to write high-performance, production-quality code.
Experience building service-oriented APIs and cloud services.
Experience designing, implementing, and deploying microservices.
Extremely technical, hands-on server software experience.
Proficient in Golang and JavaScript, with the ability to quickly learn new languages.
Experience in Linux environments and a strong understanding of Linux fundamentals and internals, including file systems, modern memory management, threads and processes, and the user-kernel space divide.
Strong understanding of large-scale distributed systems in practice, including multi-tier architectures, application security, monitoring, and storage systems.
Working knowledge of the TCP/IP stack, internet routing, and load balancing.
Grit, drive, and a deep sense of ownership.
📌 Site Reliability Engineer (SRE) - Streaming Platform (Ciudad de México)
🏢 LTIMindtree
📍 Ciudad de México