aijobs.net

Research Engineer - LLM/VLM Inference Optimization (Seed Infra)

San Jose, California, United States

USD 254K-480K Mid-level Full Time

Apply Save
Found 5h ago
Tasks
Perks/Benefits
Skills/Tech-stack

C plus plus | C# | CPU GPU | CPU/GPU profiling | CUDA | CUDA kernel development | Containerization | Conv2D | Cutlass | Efficient CUDA kernel development | FlashAttention | GEMM | GEMV | GPU Architecture | GPU Profiling | Graph Fusion | Kernel development | Low Precision | Low-precision computation | Model Parallelism | Parallel Computing | Performance Modeling | PyTorch | Python | Speculative decoding | Streaming inference | TensorFlow | TensorRT | Triton

Education

Bachelor of Engineering | Bachelor of Science

Roles

Engineer | Research Engineer

Regions

North America

Countries

Costa Rica | United States

States

California, US

Cities

San Jose, California, US

Apply Save
Language: en Views: 0 Clicks: 0 Saves: 0

Related jobs