Engineering Manager, Deep Learning Inference
Tasks
- Build optimized inference pipelines
- Develop inference framework roadmap
- Execute OSS inference framework delivery
- Guide best practices for CUDA
- Guide best practices for CUTLASS
- Guide best practices for Triton
- Implement multi GPU communications
- Lead engineering team
- Mentor and scale engineering team
- Oversee distributed inference performance
- Perform model performance tuning
- Represent team in roadmap planning
- Run profiling and optimization
Perks/Benefits
- N/A
Skills/Tech-stack
Agile | C# | C++ | CUDA | Cutlass | Distributed Systems | Multi-GPU | NCCL | NIXL | NVSHMEM | Performance Profiling | Performance optimization | PyTorch | Python | SGLang | TensorRT-LLM | Triton | VLLM
Education
Roles
Regions
Countries
States
Related jobs
-
Autonomous Vehicles | Cosmos | Data Curation | Data Generation | Deep learningSenior-level Full TimeUS, CA, Santa Clara R1d ago
-
Manager, GPU Accelerated Data Analytics USD 224K-431KAlgorithms | C# | C++ | CPU architecture | CUDASenior-level Full TimeUS, CA, Santa Clara R1d ago
-
Mid-level Full TimeUS, CA, Santa Clara R1d ago
-
Manager, Large Language Model Inference USD 184K-356KAPI Development | C++ | CUDA | GPU Architecture | PythonEquity | Flexible work options | Health insurance | Paid time offMid-level Full TimeUS, CA, Santa Clara R1d ago
-
Senior Engineering Manager, Object Storage - DGX Cloud USD 272K-431KAutomated testing | C++ | CI/CD | Call Management | Capacity PlanningSenior-level Full TimeUS, CA, Santa Clara R2d ago
-
Agile Development | Artificial Intelligence | CI/CD | Deep learning | DeploymentEquity | Health benefitsSenior-level Full TimeUS, CA, Santa Clara R9d ago