Engineering Manager, Deep Learning Inference
Tasks
- Build optimized inference pipelines
- Collaborate with compiler teams
- Collaborate with research teams
- Develop inference framework roadmap
- Enable multi GPU communication
- Execute OSS inference framework plans
- Guide CUDA best practices
- Guide CUTLASS best practices
- Guide Triton best practices
- Lead engineering team
- Mentor engineers
- Optimize distributed inference architectures
- Perform model performance tuning
- Represent team in roadmap planning
- Run profiling and optimization
- Scale engineering organization
Perks/Benefits
Skills/Tech-stack
Agile | C# | C++ | CUDA | Cutlass | Distributed Systems | FlashInfer | GPU Programming | Multi-GPU | Multi-GPU programming | NCCL | NIXL | NVSHMEM | Performance Profiling | Performance optimization | PyTorch | Python | SGLang | TensorRT-LLM | Triton | VLLM
Education
Bachelor of Engineering | Bachelor of Science | Master of Science | PhD
Regions
Countries
States
Related jobs
-
Agile Development | Aha! | Artificial Intelligence | CI/CD | ConfluenceEquity | Health benefits | Paid time offSenior-level Full TimeUS, CA, Santa Clara R2d ago
-
Senior Technical Program Manager, GenAI and Models USD 168K-258KAsynchronous RL | CI/CD | Confluence | Distributed Systems | Distributed TrainingBenefits | EquitySenior-level Full TimeUS, CA, Santa Clara R2d ago
-
Data Curation | Data Generation | Data Processing | Data Processing Pipelines | Deep learningSenior-level Full TimeUS, CA, Santa Clara R6d ago
-
Aha! | CUDA toolkit | CUDNN | Confluence | Cross-Functional CollaborationSenior-level Full TimeUS, CA, Santa Clara R7d ago