03 ago
|
Valce Talent Solutions
|
Jalisco
03 ago
Valce Talent Solutions
Jalisco
Diligent is looking for a Site Reliability Engineer
(SRE) with a strong software engineering background to lead
initiatives that ensure the reliability, scalability, and
performance of our applications and infrastructure. This role is
ideal for someone passionate about automation, observability, and
continuous improvement in production environments. The idóneo
candidate is a seasoned Site Reliability Engineer with a strong
background in automation, performance optimization, and operational
excellence. You are passionate about building resilient systems and
thrive in fast-paced, complex environments. Strong communication
skills will be required as you will be working with multiple
engineering teams to implement more efficient and reliable
procedures. Key Responsibilities • Support multiple Diligent
datacentres and deployments in the Cloud (AWS, Azure, GCP). • Lead
and improve daily operational processes across production
environments. • Contribute to incident management and response
efforts, conduct root cause analysis and lead retrospective reviews
for all Diligent applications. • Build self-service tools and
automation pipelines to enhance developer productivity and system
reliability. • Drive observability, cost optimization, and security
best practices across cloud environments. • Collaborate with
engineering and SRE teams to improve system performance,
maintainability, and scalability. • Ensure high availability and
fault tolerance of cloud services through proactive monitoring and
automation. Required Experience/Skills • 3+ years of professional
experience in Software Engineering, DevOps, or Site Reliability
Engineering. • Provide support to development teams on monitoring,
scalability, and reliability. • Intermediate-level experience with
AWS, including services like EC2, Lambda, ECS, Fargate, S3, IAM,
VPC, Route 53, RDS, DynamoDB, and CloudWatch. • Experience with
Infrastructure as Code (IaC) using Terraform, Terraform CDK or AWS
CDK. • Proficiency in designing, implementing, and testing scalable
software architectures. • Strong automation and scripting abilities
for operational workflows and cloud infrastructure. • Demonstrated
experience using AI tools to accelerate engineering workflows (code
generation, test writing, documentation, RCA analysis) •
Familiarity with LLM APIs and prompt engineering for automation
tasks • Experience with building AI-assisted runbooks, alert
triage, or auto-remediation workflows • Ability to critically
evaluate AI-generated code and configurations for correctness,
security, and operational risk • Excellent problem-solving skills.
Preferred Experience/Skills • Practical use of AI coding assistants
(GitHub Copilot, Cursor) in day-to-day engineering • Experience
with Kubernetes (EKS, Fargate, or Rancher). • Familiarity with
monitoring and observability tools like OpenSearch (Elastic),
SignalFX, or AWS-native services. • Experience with event-driven
architectures and serverless computing. • Intermediate-level
Object-Oriented Programming (OOP) skills, with experience in
TypeScript or Python. • Good working experience for a wide range of
technologies, frameworks, and platforms such as AWS Cloud, Windows,
Linux, and Kubernetes.
📌 Site Reliability Engineer (Jalisco)
🏢 Valce Talent Solutions
📍 Jalisco