Engineering Manager, Deep Learning Inference
Tasks
- Develop end to end optimized inference pipelines
- Drive inference framework strategy roadmap execution
- Ensure alignment with AI software strategy
- Foster technical excellence open collaboration continuous innovation
- Guide CUDA Triton CUTLASS adoption
- Lead engineering team for deep learning inference
- Manage multi GPU communications for inference
- Mentor and scale engineering team
- Oversee performance tuning profiling optimization of models
- Represent team in roadmap and planning discussions
Perks/Benefits
- N/A
Skills/Tech-stack
Agile | C# | C++ | CUDA | Cutlass | Deep learning | Distributed Systems | GPU Programming | LLM serving | Language Models | Large Language Models | NCCL | NIXL | NVSHMEM | Performance optimization | Profiling | Python | Triton
Education
Regions
Countries
States
Related jobs
-
Autonomous Vehicles | Cosmos | Data Curation | Data Generation | Deep learningSenior-level Full TimeUS, CA, Santa Clara R1d ago
-
Manager, GPU Accelerated Data Analytics USD 224K-431KAlgorithms | C# | C++ | CPU architecture | CUDASenior-level Full TimeUS, CA, Santa Clara R1d ago
-
Manager, Large Language Model Inference USD 184K-356KAPI Development | C++ | CUDA | GPU Architecture | PythonEquity | Flexible work options | Health insurance | Paid time offMid-level Full TimeUS, CA, Santa Clara R1d ago
-
Senior Engineering Manager, Object Storage - DGX Cloud USD 272K-431KAutomated testing | C++ | CI/CD | Call Management | Capacity PlanningSenior-level Full TimeUS, CA, Santa Clara R2d ago
-
Mid-level Full TimeUS, CA, Santa Clara R9d ago
-
Agile Development | Artificial Intelligence | CI/CD | Deep learning | DeploymentEquity | Health benefitsSenior-level Full TimeUS, CA, Santa Clara R9d ago