aijobs.net

Senior Inference Runtime Engineer

Singapore, SG / Penang, MY / Taiwan

TWD 1900K-2500K (estimate) Senior-level Full Time

Apply Save
Found 1d ago
Tasks
Perks/Benefits
Skills/Tech-stack

CUDA | CUDA profiling | Continuous batching | Distributed inference | GPU Memory Optimization | GPU memory | Go | HBM Bandwidth | KV cache | LLM serving | Memory Optimization | Model Quantization | NCCL | Pipeline parallelism | Python | Speculative decoding | Streaming | Tensor Parallelism | TensorRT-LLM | Triton | VLLM

Education

N/A

Roles

Engineer | Learning Engineer | Machine Learning Engineer | Runtime Engineer | Senior Inference Runtime Engineer | Software Engineer

Regions

Asia/Pacific

Countries

Singapore | Taiwan

Cities

Singapore, SG

Apply Save
Language: en Views: 0 Clicks: 0 Saves: 0

Related jobs