06 ago
|
BairesDev
|
México
At BairesDev®, we've been leading the way in technology projects for over 15 years. We deliver cutting-edge solutions to giants like Google and the most innovative startups in Silicon Valley.
Our diverse 4,000+ team, composed of the world's Top 1% of tech talent, works remotely on roles that drive significant impact worldwide.
When you apply for this position, you're taking the first step in a process that goes beyond the ordinary. We aim to align your passions and skills with our vacancies, setting you on a path to exceptional career development and success.
Reliability Engineer at BairesDev
As a Reliability Engineer, you will ensure the high availability and performance of production systems through engineering-led operations. You will act as a guardian of system health, balancing the need for rapid innovation with the necessity of maintaining stable and resilient infrastructure.
What You'll Do
- Establish and monitor critical reliability metrics, including SLOs, SLIs, and error budgets, to drive data-informed operational decisions.
- Build and maintain robust observability frameworks to provide deep visibility into system performance and health across complex environments.
- Orchestrate the incident lifecycle from initial detection through resolution, ensuring minimal business impact and swift service restoration.
- Drive continuous improvement by leading blameless post-mortems and performing root cause analysis to prevent the recurrence of production failures.
- Automate repetitive operational tasks and manual toil through advanced scripting and Infrastructure as Code practices.
- Conduct capacity planning and system assessments to ensure infrastructure can scale efficiently to meet evolving demand.
What We Are Looking For
- 4+ years of experience in Site Reliability Engineering, DevOps, or Software Engineering.
- Proven expertise in managing production reliability, SLOs/SLIs, and error budgets.
- Strong proficiency in Linux systems administration and scripting with Python or Go.
- Hands-on experience with Kubernetes and observability tools like Prometheus, Grafana, or Datadog.
- Demonstrated ability in incident management and at-scale operations.
- Advanced proficiency in English.
How we do make your work (and your life) easier:
- 100% remote work (from anywhere).
- Excellent compensation in USD or your local currency if preferred
- Hardware and software setup for you to work from home.
- Adaptable hours: create your own schedule.
- Paid parental leaves, vacations, and national holidays.
- Innovative and multicultural work environment: collaborate and learn from the global Top 1% of talent.
- Supportive environment with mentorship, promotions, skill development, and diverse growth opportunities.
Join a global team where your unique talents can truly thrive and make a significant impact! Apply now!
📌 Reliability Engineer - Remote Work (México)
🏢 BairesDev
📍 México