Staff Engineer, Inference Optimizations
Tasks
- Contribute to GPU infrastructure open source communities
- Develop and deploy quantization techniques
- Drive systems design for low level GPU programming
- Engineer attention layer optimizations
- Identify kernel fusion opportunities for transformer layers
- Implement advanced parallelization across multi node GPU clusters
- Lead design reviews and technical mentorship
- Lead performance architecture for inference engines
- Manage memory and precision during inference
- Optimize inference benchmarks and GPU kernel performance
- Partner with product teams on shippable performance features
- Tune mixture of experts router kernels
Perks/Benefits
- Conference reimbursement
- Education reimbursement
- Employee assistance program
- Employee stock purchase program
- Equity compensation
- Flexible time off
- LinkedIn Learning access
- Local Employee Meetups
- Remote work
- Training reimbursement
Skills/Tech-stack
Attention Optimization | BF16 | CUDA | FP4 | FP8 | GPU Architecture | INT8 | Kernel Fusion | Memory Optimization | Mixture of Experts | Parallel Computing | Precision Management | Quantization | ROCm | TensorRT | Triton
Education
N/A
Roles
Regions
Countries
States
Related jobs
-
Machine Learning Research Engineer USD 200K-330KAWS | Azure | Benchmarking | CUDA | DDP401k match | Dental insurance | Health insurance | Paid time off | Vision insuranceMid-level Full TimeEmeryville, California, United States; Hybrid (2-3 … R16h ago
-
Staff Engineer, Inference Optimizations USD 191K-239KBF16 | Bandwidth Optimization | CUDA | Expert Routing | FP4Conference reimbursement | Education reimbursement | Employee assistance program | Flexible time off | LinkedIn Learning accessSenior-level Full TimeDenver R16h ago
-
Staff Engineer, Inference Optimizations USD 191K-239KAttention Mechanisms | BF16 | CUDA | CUDA kernels | Distributed SystemsConference reimbursement | Education reimbursement | Employee assistance program | Flexible time off | LinkedIn LearningSenior-level Full TimeBoston R16h ago
-
Staff Engineer, Inference Optimizations USD 191K-239KAMD GPU | Attention Optimization | Bandwidth Optimization | Benchmarking | CUDAConference reimbursement | Employee assistance program | Flexible time off | LinkedIn Learning access | Local Employee MeetupsSenior-level Full TimeAustin R16h ago
-
Senior Software Engineer I - AI Inference Data Plane USD 139K-174KAutoscaling | Continuous batching | Data parallelism | Distributed Systems | GRPCConference reimbursement | Education reimbursement | Employee assistance program | Employee stock purchase program | Equity compensationSenior-level Full TimeAustin R17h ago
-
Senior Engineer, Inference Data Plane USD 139K-174KContinuous batching | Data parallelism | Distributed Systems | GRPC | GoEmployee assistance program | Employee stock purchase program | Equity compensation | Flexible time off | LinkedIn Learning accessSenior-level Full TimeDenver R17h ago
-
Senior Engineer, Inference Data Plane USD 139K-174KContinuous batching | Data parallelism | Distributed Systems | GRPC | GoConference reimbursement | Employee assistance program | Employee stock purchase program | Equity compensation | Flexible time offSenior-level Full TimeBoston R17h ago
-
Senior Software Engineer I - AI Inference Data Plane USD 139K-174KAutoscaling | Continuous batching | Data parallelism | Distributed Systems | GRPCConference reimbursement | Employee assistance program | Employee stock purchase program | Equity compensation | Flexible time offSenior-level Full TimeSan Francisco R17h ago
-
Senior Embedded Software Engineer - Future Forward USD 134K-201KAuthentication | C# | C++ | CAN | CUDASenior-level Full TimeSunnyvale, CA, United States R1d ago
-
Asynchronous training | Containerization | Curriculum learning | Distributed Systems | Distributed TrainingSenior-level Full TimeSF Bay Area, CA, Remote, US, … R2d ago
-
Principal Machine Learning Engineer, Content Safety USD 295K-345KComputer Vision | Content Moderation | Data Pipelines | Deep learning | Language ModelsEquity compensationSenior-level Full TimeSan Mateo, CA, United States R3d ago
-
Machine Learning Infrastructure Engineer USD 100K-150KAPI Gateway | Abuse detection | Automated rollback | Autoscaling | C++Senior-level Full TimeUnited States - Remote R3d ago
-
Senior-level Full TimeUnited States - Remote R3d ago
-
LLM Engineer USD 100K-150KAdapter based Fine Tuning | Attention Optimization | Cluster management | DPO | Distributed TrainingSenior-level Full TimeUnited States - Remote R3d ago
-
AI Optimization Engineer USD 100K-150KBenchmarking | C++ | CUDA | Continuous batching | CutlassRemote workSenior-level Full TimeUnited States - Remote R3d ago
-
Data Engineer, Principal (Hybrid) USD 170K-190KAgile | Azure Synapse | Azure Synapse Analytics | Bitbucket | CI/CDContract-to-hire opportunity | Hybrid work scheduleSenior-level ContractSan Francisco, CA R3d ago
-
AWS | AWS CDK | AWS CodeBuild | AWS CodePipeline | AWS Secrets401k | Healthcare coverage | PTO | Paid Company Holidays | Phone stipendSenior-level Full TimeSan Carlos - Hybrid R3d ago
-
Edge AI Engineer USD 100K-150KC++ | Core ML | DSP | Embedded Systems | Energy optimizationCareer growth | Remote workSenior-level Full TimeUnited States - Remote R4d ago
-
Senior-level Full TimeUnited States - Remote R4d ago
-
Senior-level Full TimeUnited States - Remote R4d ago
-
LLM Engineer USD 100K-150KAdapter modules | Attention Optimization | Benchmarking | DPO | Dataset DistillationSenior-level Full TimeUnited States - Remote R4d ago
-
LLM Engineer USD 100K-150KAdapter Method | Attention Optimization | DPO | Distributed Training | Efficient Fine TuningSenior-level Full TimeUnited States - Remote R4d ago
-
Senior-level Full TimeUnited States - Remote R4d ago
-
Model Optimization Engineer USD 100K-150KC++ | CUDA | Continuous batching | Deep learning | DeepSpeedSenior-level Full TimeUnited States - Remote R4d ago
-
ML Systems Engineer USD 100K-150KAPI Gateways | Abuse detection | Automated rollback | Autoscaling | BatchingSenior-level Full TimeUnited States - Remote R4d ago