13 ago
|
GE Vernova
|
Santiago de Querétaro
13 ago
GE Vernova
Santiago de Querétaro
The Production DevOps Engineer serves as a critical link in the "Middle-Mile" of software delivery for the GE Vernova’s Grid Software SaaS products. This role is responsible for ensuring that software moves from development to production environments through a standardized, secure, and highly observable path. You will own the Change Management Process, serving as a primary authority for production deployments to ensure that new SaaS product versions do not compromise the stability of global energy grid operations. This position requires a strong technical background in automation and a disciplined approach to release safety in a 24/7 operational environment.
Works independently and is seen as a Technical Leader. The role demonstrates deep understanding of concurrent software development, its effect on build management and releasing the builds across versions and environments
**Roles and Responsibilities**
**Day 0: Pipeline Implementation & Standardization**
- **Golden Path Execution**: Maintain and improve standardized CI/CD pipelines using GitHub Actions and ArgoCD, ensuring all product teams follow the established "Golden Path" to avoid bespoke, non-standard deployment utilities.
- **Policy Enforcement**: Implement and manage automated "quality gates" within the delivery pipeline to verify that every release meets security and architectural standards before reaching production.
- **Provisioning Support**: Assist the SaaS Cloud Engineers in automating highly secure, resilient customer’s cloud infrastructure.
**Day 1: Release Authority & Deployment Management**
- **Change Control Authority**:**Review and provide final approval for production deployment requests, ensuring all pre-release criteria—such as performance testing and security scanning—are satisfied.
- **Progressive Delivery**:**Execute advanced rollout strategies, including Canary and Blue/Green deployments on Kubernetes, to minimize the "blast radius" of changes.
- **Validation**:**Perform automated verification and acceptance testing post-deployment to confirm service health and trigger automated rollbacks if necessary.
**Day 2: Operational Support & Optimization**
- 24/7 Follow-the-Sun Support: Participate in general on-call rotations, ensuring a seamless transition of operational responsibility between time zones through standardized handover protocols.
- **Incident & Root Cause Analysis**: Support high-severity incident response and participate in blameless Root Cause Analysis (RCA) to identify and fix systemic deployment risks.
- **FinOps & Capacity**: Track and report on cloud resource consumption for CI/CD infrastructure, assisting in cost-optimization efforts and right-sizing production workloads.
**Technical Requirements**
- **CI/CD & GitOps**: Hands-on experience with Jenkins, Artifactory, GitHub Actions and ArgoCD for automated software delivery.
- **Container Orchestration**: Proficiency in managing workloads on Kubernetes, specifically with EKS clusters.
- **Automation Tools**:**Strong skills in Ansible and Terraform for configuration management and infrastructure-as-code.
- **Cloud Platform**: Solid understanding of AWS cloud services (VPC, IAM, EKS, RDS, S3, MSK, etc) in a production setting.
- **Observability**:**Experience using Prometheus,
Grafana, Splunk, Datadog or Dynatrace to monitor deployment health and system performance.
- **Scripting**:Experience building dynamic build pipelines using Groovy Script, Python, Bash or Go languages.
**Experience & Qualifications**
- **Operational Discipline**:**Proven ability to manage production changes and troubleshooting under pressure in a high-stakes environment.
- **Compliance Awareness**: Familiarity with regulated industries and security frameworks such as NERC CIP, SOC2, ISO 27001, IEC 62443 is highly preferred.
- **Communication**: Strong ability to document technical procedures and communicate clearly with stakeholders during global shift handovers.
**Education Qualification**
Bachelor's Degree in Computer Science or “STEM” Majors (Science, Technology, Engineering and Math) with advanced experience.
**Key Performance Indicators (KPIs)**
- **Customer Onboarding Speed**: Contribution towards the 4-hour SLA target.
- **System Availability**:**Help maintain 99.99% availability of mission critical grid SaaS products.
- **Change Failure Rate**:**Maintaining a low rate of failed production deployments through improved quality gates.
- **Mean Time to Recover (MTTR)**: Ensuring fast restoration of service through automated rollbacks and clear runbooks.
- **Toil Reduction**: Automating repetitive manual tasks to ensure at least 50% of time is spent on engineering improvements.
**Business Acumen**:
- Strong problem solving abilities and capable of articulating specific technical topics or assignments
- Experience in building scalable and highly available distributed systems
- Skilled in breaking down problems and estimate time for development tasks
- Evangelizes how our technology
📌 Sre Production DevOps (Santiago de Querétaro)
🏢 GE Vernova
📍 Santiago de Querétaro