aijobs.net

Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start

San Jose, California, United States

USD 128K-256K Entry-level Full Time

Apply Save
Found 5h ago
Tasks
Perks/Benefits
Skills/Tech-stack

Access Optimization | C# | C++ | CUDA | Computational Graph Optimization | Computational graph | Constant folding | Deep learning | Distributed Systems | GPU | GPU Performance | GPU memory | GPU performance analysis | Graph optimization | Memory Reuse | Memory access | Memory access optimization | Mixture of Experts | NPU | NVIDIA Nsight | Operator fusion | Performance Analysis | Pipeline parallelism | Profiling | Python | Quantization | Scheduling optimization | Sequence parallelism | Tensor Parallelism | Vectorization

Education

Bachelor of Science | Master of Science

Roles

Backend | Backend Engineer | Engineer | Learning Engineer | Machine Learning Engineer

Regions

North America

Countries

Costa Rica | United States

States

California, US

Cities

San Jose, California, US

Apply Save
Language: en Views: 1 Clicks: 0 Saves: 0

Related jobs