MLOps Engineer - GPU Platform (Inference & Training)
Almaty, Kazakhstan
R
USD 184K-300K (estimate) Senior-level Full Time
Tasks
- Automate deployments with GitOps using ArgoCD and Helm
- Collaborate on training runs and model serving rollouts
- Implement GPU aware scheduling and warm pools
- Improve GPU utilization and reduce idle allocations
- Manage Talos Linux and Sidero Omni cluster lifecycle
- Monitor GPU health and resolve node incidents
- Operate NVIDIA GPU drivers and GPU Operator
- Optimize distributed training clusters
- Provision infrastructure with Terraform
- Scale inference with KEDA and Kafka autoscaling
- Set observability with Prometheus and VictoriaMetrics
- Track SLOs and GPU cost per generation
- Tune NCCL and GPU network performance
Perks/Benefits
Skills/Tech-stack
Amazon EKS | ArgoCD | Bottlerocket | CUDA | DCGM | GPU Operator | GVisor | GitOps | Go | Grafana | Helm | Infiniband | Istio | KEDA | Kafka | Karpenter | Kubernetes | Linux | Linux Networking | NCCL | NVIDIA GPU | NVIDIA GPU Operator | Nebius | Prometheus | Python | RoCE | Sidero Omni | Talos Linux | Terraform | Terraform Cloud | VictoriaMetrics
Education
N/A
Roles
Related jobs
- No jobs found.