04 ago
|
Luxoft
|
Guadalajara
04 ago
Luxoft
Guadalajara
- Responsibilities:
- Operate and support Linux servers and shared infrastructure used by research and platform teams.
- Troubleshoot production issues involving system performance, availability, access, configuration, and networking.
- Support Kubernetes and other container-based platforms, including node health, service behavior, and rollout activities.
- Support scheduler-backed compute environments such as Slurm, including node readiness, maintenance, and incident recovery.
- Improve user-facing Linux access services such as SSH, shared shell environments, and session-based platforms.
- Manage OS lifecycle work: provisioning, patching, kernel and package updates, and hardening.
- Build scripts and automation in Python, Bash, or similar tools to reduce manual work and improve reliability.
- Use configuration management and version-controlled workflows to implement infrastructure changes safely.
- Enhance monitoring, alerting, documentation, and operational processes.
- Participate in incident response and occasional on-call support.
- Mandatory Skills Description:
- 3+ years of experience in Linux systems engineering, SRE, DevOps, or infrastructure support.
- Strong Linux administration skills, including systemd, package management, permissions, filesystems, log analysis, and performance troubleshooting.
- Good understanding of networking fundamentals such as DNS, NTP/PTP, routing, and general host connectivity.
- Experience with automation, scripting, and operational tooling.
- Familiarity with Kubernetes, virtualization, or clustered platforms.
- Experience with configuration management or infrastructure-as-code tools such as Ansible, Salt, or Terraform.
- Ability to troubleshoot production issues methodically and communicate clearly during incidents.
- Experience with Git-based workflows and maintainable documentation.
- Hands-on, practical problem solver with a strong ownership mindset.
- Comfortable working close to production and balancing support with continuous improvement.
- Collaborative communicator who works well across compute, storage, networking, and application teams.
- Nice-to-Have Skills Description:
- Experience with Slurm, HPC-style environments, GPU infrastructure, or researcher-facing Linux platforms.
- Working familiarity with shared storage clients such as NFS, autofs, or GPFS / IBM Storage Scale from a host and application perspective.
- Experience with observability tools such as Prometheus, Grafana, or equivalent platforms.
- Exposure to identity and access services such as LDAP, Kerberos, SSSD, or PAM.
- Exposure to on-premises datacenter operations, hardware lifecycle support, or vendor escalations.
- Interest in using AI/ML techniques for infrastructure optimization, anomaly detection, or predictive operations.
📌 Linux Systems Engineer (Guadalajara)
🏢 Luxoft
📍 Guadalajara