Manager, Large Language Model Inference
Tasks
- Coordinate cross-functional teams
- Design production inference software
- Develop LLM inference runtime software
- Integrate NVIDIA technologies
- Lead engineering team
- Optimize inference kernels
- Plan projects and deliver milestones
- Provide developer experience for LLM deployment
Perks/Benefits
Skills/Tech-stack
API Development | C++ | CUDA | GPU Architecture | Python | SGLang | TensorRT | TensorRT-LLM | VLLM
Education
Roles
Regions
Countries
States
Related jobs
-
Manager, GPU Accelerated Data Analytics USD 224K-431KAlgorithms | C# | C++ | CPU architecture | CUDASenior-level Full TimeUS, CA, Santa Clara R1d ago
-
Mid-level Full TimeUS, CA, Santa Clara R1d ago
-
Senior Engineering Manager, Object Storage - DGX Cloud USD 272K-431KAutomated testing | C++ | CI/CD | Call Management | Capacity PlanningSenior-level Full TimeUS, CA, Santa Clara R2d ago
-
Mid-level Full TimeUS, CA, Santa Clara R9d ago