Design and operate scalable infrastructure for training and serving large ML workloads, focusing on cost-efficiency, reliability and reproducibility.
Tasks
Architect and maintain scalable GPU/TPU clusters and provisioning workflows.
Implement data storage, versioning and feature store infrastructure.
Automate capacity planning, cost controls and job scheduling for training workloads.
Collaborate on security, multi-tenancy and access patterns for research infra.
Requirements
4+ years in infra engineering with experience in cloud GPU/accelerator environments.
Experience with cluster orchestration, provisioning and performance tuning.
Familiarity with storage systems, feature stores and large-data workflows.
Strong scripting and automation skills (Python, Bash, Terraform).