26 ago
|
Agileengine
|
Xico
Job Description
Agile Engine is an Inc. **** company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries.
We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.WHY JOIN US
If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!
ABOUT THE ROLE
We are looking for anSRE Operations Engineerto keep production and staging environments running reliably across a cloud-based Saa S platform.
You'll respond to live incidents, reduce operational toil through automation, and improve observability using Kubernetes, Terraform, Grafana, and AWS.
A hands-on role with real ownership across CI/CD pipelines, Git Ops workflows, and on-call rotations.WHAT YOU WILL DO
- Monitor and support production and staging environments in real time, ensuring high availability, performance, and stability;
- Respond to incidents, perform triage and root cause analysis, and contribute to post-incident reviews and remediation efforts;
- Participate in an on-call rotation with defined SLAs;
- Handle ad-hoc and unplanned operational requests from Product, Support, and internal teams;
- Maintain and enhance monitoring, alerting, dashboards, logs, and metrics, and improve observability practices;
- Support CI/CD pipelines, production releases, and Git Ops workflows;
- Contribute to automation efforts to reduce operational toil;
- Maintain and improve Kubernetes-based infrastructure and containerized workloads;
- Support Infrastructure as Code practices and ongoing environment improvements.MUST HAVES
-2+ years of experienceinSite Reliability Engineering ,Dev Ops , orProduction Operations ;
- Experience withAWSsupporting production environments;
- Experience supportingproduction Saa S applications ;
- Strong understanding ofCI/CD systemssuch asGit Hub Actions ,Jenkins , orCircle CI ;
- Experience withGit Opsand strongGitfundamentals;
- Experience usingGit Hub ,Jira , andConfluencein collaborative environments;
- Experience withKubernetessuch asEKSork Ops ;
- Experience withDockerand containerization;
- Experience withobservability toolssuch asGrafana ,Prometheus ,Loki , orPager Duty ;
- Experience withscripting languagessuch asBash ,Python , orGo ;
- Experience withInfrastructure as Codesuch asTerraformorHelm ;
- Ability to work within structured operational processes and SLAs;
- Strong written and verbal English communication skills;
- Self-driven with a growth mindset.NICE TO HAVES
- AWS certifications such as Solutions Architect, Dev Ops Engineer, or Sys Ops Administrator;
- Experience in multi-tenant Saa S environments;
- Experience working in globally distributed teams;
- Familiarity with Chat Ops practices;
- Experience improving monitoring quality and reducing alert fatigue.PERKS AND BENEFITS
-Professional growth:Mentorship, Tech Talks, and personalized growth roadmaps.
-Competitive compensation:USD-based pay with education, fitness, and team activity budgets.
-Exciting projects:Modern solutions with Fortune 500 and top product companies.
-Flextime:Adaptable schedule with remote and office options.
📌 Site Reliability Engineer Id60188 (Xico)
🏢 Agileengine
📍 Xico