Manager, Large Language Model Inference
Tasks
- Coordinate cross-functional teams
- Design production inference software
- Develop LLM inference runtime software
- Integrate NVIDIA technologies
- Lead engineering team
- Optimize inference kernels
- Plan projects and deliver milestones
- Provide developer experience for LLM deployment
Perks/Benefits
Skills/Tech-stack
API Development | C++ | CUDA | GPU Architecture | Python | SGLang | TensorRT | TensorRT-LLM | VLLM
Education
Roles
Regions
Countries
States
Related jobs
-
Software Engineering Manager - Cloud Software, Data Platform & AI-Agentic Product Engineering USD 200K-215KAI | Call Management | Cloud Computing | Data platforms | Development Lifecycle401k | Dental insurance | Health insurance | Health savings account | Life insuranceMid-level Full TimeSanta Clara, CA R4d ago
-
Manager, GPU Accelerated Data Analytics USD 224K-431KAlgorithms | C# | C++ | CPU architecture | CUDASenior-level Full TimeUS, CA, Santa Clara R21d ago
-
Mid-level Full TimeUS, CA, Santa Clara R21d ago
-
Senior Engineering Manager, Object Storage - DGX Cloud USD 272K-431KAutomated testing | C++ | CI/CD | Call Management | Capacity PlanningSenior-level Full TimeUS, CA, Santa Clara R22d ago
-
Mid-level Full TimeUS, CA, Santa Clara R29d ago