aijobs.net

Software Engineer, GPU Inference

US and Canada Offices

CAD 135K-220K (estimate) Senior-level Full Time

Apply Save
Found 2d ago
Tasks
Perks/Benefits
Skills/Tech-stack

Asynchronous execution | BF16 | Benchmarking automation | Benchmarks | C++ | CI/CD | Collective communication | Concurrency | Containers | Continuous batching | Distributed Systems | Expert parallelism | FP4 | FP8 | GPU | HIP | INT4) | INT8 | Inference Server | KV cache | Kernel launch | Kubernetes | Linux | Memory Management | Mixture of Experts | Multithreading | Numerical validation | Observability | Pipeline parallelism | Prefix caching | Profiling | PyTorch | Python | Quantization | RDMA | ROCm | SGLang | Synchronization | Tensor Parallelism | TensorRT-LLM | Triton Inference | Triton Inference Server | VLLM

Education

Bachelor of Engineering | Bachelor of Science | Bachelor of Science in Computer Science

Roles

Engineer | Software Engineer

Regions

North America

Countries

Canada

Apply Save
Language: en Views: 0 Clicks: 0 Saves: 0

Related jobs