19 ago
|
Net2Source
|
México
job title: site reliability engineer (sre) - streaming platform
work mode: remote
job description:
this is a critical role with a wide range of responsibilities, including:
- analyze and improve system design to reduce failure modes and promote self-healing systems.
- establish and maintain robust systems that facilitate observability, encompassing logging, monitoring, distributed tracing, alerting, and offline test tools.
- work with development partners to shape the architecture, design, and implementation of new and existing systems to enhance their reliability, performance, efficiency, and scalability.
- ability to work both independently and as part of a geographically dispersed, yet integrated, team.
- collaborate with service engineers to establish service level agreements (slas) and service level objectives (slos) for backend services.
- identify indications or cues that demonstrate the effectiveness of an application and possess the knowledge to improve or repair its performance.
- assess options and suggest solutions when information is limited or unclear. This position requires a level of comfort and confidence in dealing with uncertain situations.
- work seamlessly within a team while effectively managing individual tasks.
- respond to emerging incidents, solve critical issues, and follow through with a plan for resolution or future mitigation.
- act as an sme on the engineering operations team,
partnering with backend services teams and application teams to overcome challenges across all platforms where we stream our service.
qualities & experience we're seeking
we believe the right individual will have the following skills and experience to be successful in the role:
- 5+ years of experience in software development.
- degree in computer science or a related field, or equivalent work experience.
- solid engineering and coding skills, strong data structure knowledge, and the ability to write high-performance, production-quality code.
- experience building service-oriented apis and cloud services.
- experience designing, implementing, and deploying microservices.
- extremely technical, hands-on server software experience.
- proficient in golang and javascript, with the ability to quickly learn new languages.
- experience in linux environments and a strong understanding of linux fundamentals and internals, including file systems, modern memory management, threads and processes, and the user-kernel space divide.
- strong understanding of large-scale distributed systems in practice, including multi-tier architectures, application security, monitoring, and storage systems.
- working knowledge of the tcp/ip stack, internet routing, and load balancing.
- grit, drive, and a deep sense of ownership.
📌 Site reliability engineer (sre) - streaming platform (México)
🏢 Net2Source
📍 México