AgileEngine is seeking a Senior Site Reliability Engineer to provide core system administration and operational stability for enterprise on-premise and SaaS-hosted systems. The role emphasizes Kubernetes cluster management, monitoring, and observability using Snowflake and OpenTelemetry.
You will participate in on-call rotations, incident response, and RCA, and you will automate infrastructure tasks with Python, Bash, or Go to ensure health across ESM and ECP environments.