Data Center Operations System Engineer III at lambda.
About the role Lambda builds the infrastructure powering the next generation of artificial intelligence. We are looking for a hands-on engineer to manage our physical data center operations in Mexico. This role requires daily on-site presence and shift work to maintain our high-performance GPU clusters.
Key facts
- Location: Querétaro, Mexico
- Engagement: Full-time, On-site
- Compensation: MX$531K - MX$708K
- Team: Data Center Business
What you'll do
- Install, rack, cable, and configure server, storage, and networking hardware.
- Diagnose and resolve hardware and software failures within advanced GPU and network environments.
- Maintain accurate records of network topology and data center layouts using DCIM tools.
- Coordinate with supply chain and manufacturing departments to execute large-scale deployments.
- Oversee parts inventory and track equipment lifecycle from delivery through final handoff.
- Collaborate with hardware support teams to document and solve complex infrastructure incidents.
- Manage the RMA process for faulty components and ensure timely replacements.
- Enforce strict installation standards to ensure consistency across all sites.
Requirements
- Proficiency in both English and Spanish.
- Experience with critical data center infrastructure, including power distribution, airflow management, environmental monitoring,
and structured cabling.
- Knowledge of three-phase and single-phase power, PDU balancing, and thermal containment strategies.
- Understanding of server hardware, boot sequences, and cable optics.
- Ability to perform carrier DIA circuit testing and fiber troubleshooting.
- Experience creating and refining maintenance procedures.
- Willingness to travel to support the launch of new data center sites.
- Ability to mentor junior staff members.
Nice to have
- 3+ years of experience with data center critical infrastructure systems.
- Familiarity with 400gb Infiniband architectures and network configurations.
- Experience with DDP or SCM cluster storage.
- 3+ years of experience using ticketing platforms like JIRA or Zendesk.
- Advanced Linux administration skills.
- Experience with high-performance GPU systems, specifically Nvidia NVL72.
Skills & tools
- DCIM software
- Linux administration
- Fiber and copper cabling
- JIRA/Zendesk
- Nvidia NVL72
- Infiniband
Practical notes
This is a non-exempt position eligible for overtime pay. You must be available to work on-site in Querétaro five days per week on a shift-based schedule. Lambda provides health, dental, and vision coverage for employees and their dependents, along with a versátil time-off policy.
#J-18808-Ljbffr
📌 Data Center Operations System Engineer III (Santiago de Querétaro)
🏢 Lambda
📍 Santiago de Querétaro
Postulate a este anuncio
Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.