aijobs.net

Staff Software Engineer, GPU Inference

Toronto Office

USD 175K-260K (estimate) Senior-level Full Time

Apply Save
Found 10d ago
Tasks
Perks/Benefits
Skills/Tech-stack

Asynchronous execution | Benchmarking | C++ | CI/CD | CUDA | Cache Management | Collective communication | Concurrency | Continuous batching | Determinism | Distributed Systems | Distributed systems debugging | Docker | Expert parallelism | Failure recovery | GPU Performance | GPU Synchronization | GPU performance profiling | HIP | KV cache | KV-cache management | Kernel Launches | Kubernetes | Linux | Memory Management | Multithreading | Numerical validation | Observability | Performance Profiling | Pipeline parallelism | Prefix caching | Profiling | PyTorch | Python | Quantization | RDMA | ROCm | Scheduling | Systems debugging | Tensor Parallelism | VLLM

Education

Bachelor of Engineering | Bachelor of Science

Roles

Engineer | Software Engineer | Staff Software Engineer

Regions

North America

Countries

Canada | United States

States

Ontario, CA | California, US

Cities

Toronto, Ontario, CA | Sunnyvale, California, US

Apply Save
Language: en Views: 0 Clicks: 0 Saves: 0

Related jobs