Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Tasks
- Automate infrastructure with IaC
- Build CI CD pipelines for infrastructure
- Design and scale storage solutions
- Evaluate and deploy new networking storage and container technologies
- Maintain observability and monitoring stacks
- Manage and optimize job scheduling with Slurm
- Mentor engineers and enforce engineering standards
- Own GPU compute cluster lifecycle
- Tune cluster configuration for distributed training workloads
Perks/Benefits
- N/A
Skills/Tech-stack
Ansible | Apptainer | Bash | CI/CD | Cgroups | Configuration Management | Container Runtime | DCGM | Docker | EFA | File System | GitLab | Grafana | Infiniband | Kubernetes | Linux | Lustre | MIG | NFS | NVIDIA GPU | NVLink | NVSwitch | Observability | Parallel file system | Prometheus | Python | RDMA | RoCE | Singularity | Slurm | Terraform | WekaFS
Education
Related jobs
- No jobs found.