06 ago
|
BairesDev
|
México
At BairesDev®, we've been leading the way in technology projects for over 15 years. We deliver cutting-edge solutions to giants like Google and the most innovative startups in Silicon Valley. Our diverse 4,000+ team, composed of the world's Top 1% of tech talent, works remotely on roles that drive significant impact worldwide.
When you apply for this position, you're taking the first step in a process that goes beyond the ordinary. We aim to align your passions and skills with our vacancies, setting you on a path to exceptional career development and success.
Reliability
Engineer at BairesDev As a Reliability Engineer, you will ensure the high availability and performance of production systems through engineering-led operations. You will act as a guardian of system health, balancing the need for rapid innovation with the necessity of maintaining stable and resilient infrastructure.
What You'll Do Establish and monitor critical reliability metrics, including SLOs, SLIs, and error budgets, to drive data-informed operational decisions.
Build and maintain robust observability frameworks to provide deep visibility into system performance and health across complex environments.
Orchestrate the incident lifecycle from initial detection through resolution, ensuring minimal business impact and swift service restoration.
Drive continuous improvement by leading blameless post-mortems and performing root cause analysis to prevent the recurrence of production failures.
Automate repetitive operational tasks and manual toil through advanced scripting and Infrastructure as Code practices.
Conduct capacity planning and system assessments to ensure infrastructure can scale efficiently to meet evolving demand.
What We Are Looking For 4+ years of experience in Site Reliability Engineering, DevOps, or Software Engineering.
Proven expertise in managing production reliability, SLOs/SLIs, and error budgets.
Strong proficiency in Linux systems administration and scripting with Python or Go.
Hands-on experience with Kubernetes and observability tools like Prometheus, Grafana, or Datadog.
Demonstrated ability in incident management and at-scale operations.
Advanced proficiency in English. How we do make your work (and your life) easier: 100% remote work (from anywhere).
Excellent compensation in USD or your local currency if preferred
Hardware and software setup for you to work from home.
Flexible hours: create your own schedule.
Paid parental leaves, vacations, and national holidays.
Innovative and multicultural work environment: collaborate and learn from the global Top 1% of talent.
Supportive environment with mentorship, promotions, skill development, and diverse growth opportunities. Join a integral team where your unique talents can truly thrive and make a significant impact! Apply now!
📌 Reliability Engineer - Remote Work (México)
🏢 BairesDev
📍 México