Find jobs in AI/ML, Data Science and Big Data
55 results
for Speculative decoding
(Skill/Tech stack)
-
Senior Inference Runtime Engineer TWD 1900K-2500KCUDA | CUDA profiling | Continuous batching | Distributed inference | GPU Memory OptimizationFlexible work culture | Inclusive environment | Training and mentoringSenior-level Full TimeSingapore, SG / Penang, MY / …1d ago
-
Intern - GenAI Benchmarking (MLE) / On-Device Model Deployment (SWE) KRW 45000K-55000KAIMet | Automated testing | Cache Management | Code review | Context modelingEntry-level InternshipSeoul, Korea, Republic of1d ago
-
Model Optimization Engineer USD 100K-150KC++ | CUDA | Continuous batching | Deep learning | DeepSpeedSenior-level Full TimeUnited States - Remote R1d ago
-
ML Performance Engineer USD 100K-150KBenchmarking | C++ | Continuous batching | Cutlass | Deep learningCareer growth | Direct W2 employment | Remote workSenior-level Full TimeTempe, AZ R1d ago
-
AI Inference Engineer GBP 80K-106KAutoscaling | Batching | CUDA | Cache Management | Capacity PlanningBiannual bonus | Breakfast allowance | Dinner allowance | Equity sign-on bonus | Expensed technologyMid-level Full TimeLondon, England, United Kingdom - Remote R1d ago
-
AI Optimization Engineer USD 100K-150KBenchmarking | C++ | Cache optimization | Compiler optimization | Continuous batchingCareer growthSenior-level Full TimeUnited States - Remote R4d ago
-
Senior Director, Inference Products and Optimizations USD 274K-343KAI workload | AI workload orchestration | AMD GPU | CUDA | Container RuntimeConference reimbursement | Education reimbursement | Employee assistance program | Employee stock purchase program | Equity compensationSenior-level Full TimeSeattle4d ago
-
Senior Director, Inference Products and Optimizations USD 274K-343KAI workload | AI workload orchestration | AMD | Benchmarking | Call ManagementEmployee assistance program | Flexible time off | LinkedIn LearningSenior-level Full TimeSan Francisco R4d ago
-
Senior-level Full TimeUnited States - Remote R5d ago
-
Senior Software Engineer - AI Inference USD 152K-287KAttention | Batching | C++ | CUDA | ConcurrencySenior-level Full TimeUS, CA, Santa Clara R5d ago
-
Agentic AI Engineer (LLM) USD 106K-181KA/B | A/B Testing | API Development | Agent Orchestration | AutogenAnnual health check-ups | Opportunity to collaborate and learn from industry professionals | Performance bonuses | Preferential pricing for services | Premium healthcare packageMid-level Full TimeHanoi, Vietnam5d ago
-
Staff Software Engineer, Beyond Live, DeepMind USD 207K-301K2D Games | 3D Games | Audio Tokenization | Audio video pacing | Audio/VideoSenior-level Full TimeNew York, NY, USA; Mountain View, …6d ago
-
Senior Staff Engineer, GDC AI Inference Platform USD 262K-365KAgent coordination | C++ | Cloud infrastructure | Containerization | Data StructuresBonus | Equity | Health insurance | Paid time off | Retirement plansSenior-level Full TimeSunnyvale, CA, USA; Kirkland, WA, USA6d ago
-
Senior-level Full TimeCupertino7d ago
-
C plus plus | CI/CD | Cloud Computing | Computer Vision | Deep learningSenior-level Full TimeBurlingame, CA10d ago
-
AI Engineer (Managed Services) SGD 85K-138KARES | AWQ | Agent Orchestration | Agent systems | Attention MechanismsMid-level Full TimeSingapore11d ago
-
AI Engineer – Enterprise, Data & AI USD 160K-246KAWS Bedrock | Agentic Workflows | Alerting | Autogen | Autonomous AgentsMid-level Full TimeFoster City, CA11d ago
-
Machine Learning Engineer - Model Inference USD 166K-230KC++ | Cloud infrastructure | Containers | Continuous batching | Distributed SystemsMid-level Full TimeCupertino11d ago
-
Entry-level Full Time北京12d ago
-
Mid-level Full TimeSingapore12d ago
-
Tech Lead Manager, Inference USD 207K-300KAutoscaling | Cache Management | Caching | Continuous batching | Deployment PipelinesSenior-level Full TimeSF Bay Area, CA13d ago
-
Software Engineer - Training/Inference (C++) USD 180K-440KAuto Scaling | C++ | CI/CD | CUDA | Code generation401k plan | Dental insurance | Disability insurance | Discounts | Health insuranceSenior-level Full TimePalo Alto, CA14d ago
-
AWS | Alertmanager | Amazon EKS | Amazon SageMaker | AutoscalingSenior-level Full TimeGLASGOW, LANARKSHIRE, United Kingdom15d ago
-
Senior-level Full TimePangyo (Software Dream Center), South Korea15d ago
-
Staff ML Engineer, Fine Tuning - Slack USD 197K-344KDeep learning | GPU infrastructure | Go | Hybrid Retrieval Generation | Hybrid retrieval401k | Dental insurance | Employee stock purchasing program | Life and disability insurance | Medical insuranceSenior-level Full TimeWashington - Seattle, United States15d ago
-
Member of Technical Staff - Inference Research USD 150K-350KAutoscaling | Disaggregated Prefill | FP8 | INT4) | KV cacheIn-person collaborationSenior-level Full TimeNew York15d ago
-
Senior Software Engineer, Inference USD 152K-204KBF16 | C++ | CI/CD | CUDA | CUDA kernels401k employer match | Company paid life insurance | Employee stock purchase program | Flexible PTO | Flexible spending accountSenior-level Full TimeSunnyvale, CA / Bellevue, WA19d ago
-
Senior ML Engineer (Token Factory) GBP 80K-130KCI/CD | Distributed Training | Inference Optimization | JAX | JAX Speculative DecodingCareer growth and learning opportunities | Collaborative culture | Flexibility | International environment | OwnershipSenior-level Full TimeGermany; Israel; Netherlands; Prague, Czech Republic; … R20d ago
-
Principal LLM Inference Engineer USD 195K-285KBatching | C# | C++ | CUDA | CUDA kernelEquity | Flexible working hours | Health insurance | Paid time offSenior-level Full TimeSanta Clara20d ago
-
[2026] Senior Machine Learning Engineer (Systems), Embodied AI/NPCs, ML Platform - PhD Early Career USD 196K-243KAWS | Azure | Cloud platform | Continuous batching | Data PipelinesEquity compensation | Health benefits | Paid time offSenior-level Full TimeSan Mateo, CA, United States R20d ago
-
[2026] Senior Machine Learning Engineer (Systems), Embodied AI/NPCs, ML Platform - PhD Early Career USD 196K-243KAWS | Azure | Cloud platform | Continuous batching | Deep learningSenior-level Full TimeSan Mateo, CA, United States R20d ago
-
Senior Forward Deployed Engineer II (AI/ML) INR 1800K-3500KAgents SDK | CUDA | Cache optimization | Continuous batching | CrewAIMid-level Full TimeBengaluru24d ago
-
Senior Forward Deployed Engineer I (AI/ML) INR 3000K-4800KAgents SDK | CUDA | Continuous batching | CrewAI | Data CompressionHybrid work | Travel up to 30%Senior-level Full TimeBengaluru24d ago
-
Engineering Manager, ML Performance USD 207K-301KAuto sharding | Benchmarking | CUDA | CUDA Performance | Compiler optimizationSenior-level Full TimeSunnyvale, CA, USA; Kirkland, WA, USA26d ago
-
Applied AI Scientist - On Site EUR 54K-86KC++ | CUDA | Computer Vision | Deep learning | Distributed TrainingOn-site workSenior-level Full TimeMünchen, BY, DE27d ago
-
C++ | Computer Vision | Deep learning | Distributed Training | Efficient InferenceCore research and development team | On-site workSenior-level Full TimeTel Aviv-Yafo, Tel Aviv District, IL27d ago
-
Senior-level Full Time上海、北京28d ago
-
Continuous batching | Data parallelism | Deep learning | Distributed Training | Dynamic MemoryComputational resources access | Full sponsorship | Hired by Rakuten Asia after completion | Research exchangesMid-level Full TimeCrimson House Singapore1mo ago
-
Application Software Engineer, Inference USD 135K-185KAgent Orchestration | Agent SDK | Auto Scaling | Batch scheduling | C++401k plan | Employee stock purchase plan | Long-term incentives | Medical, dental & vision coverage | Onsite Palo AltoEntry-level Full TimePalo Alto, CA1mo ago
-
Sr GenAI Infra Specialist SA, AWS WWSO Startup USD 153K-228KAWS Inferentia | AWS Trainium | Amazon Web Services | Batching | CUDASenior-level Full TimeNew York, New York, USA1mo ago
-
AI Engineer EUR 60K-80KAWQ | AWS | Agent SDK | CI/CD | CUDACareer growth opportunities | Permanent employment | Remote work optionMid-level Full TimeRemote - Paris, France R1mo ago
-
Senior Machine Learning Engineer USD 188K-282KAdversarial Training | Calibration monitoring | Continuous batching | DPO | Deep learningSenior-level Full TimePalo Alto, CA1mo ago
-
Sr. Software Engineer, Inference PLN 321K-470KAutoscaling | BF16 | C++ | CI/CD | CUDACritical illness cover | Employee assistance programme | Family dental insurance | Family medical insurance | Life assuranceSenior-level Full TimeWarsaw, Poland1mo ago
-
LLM Inference Frameworks and Optimization Engineer USD 160K-230KC++ | CUDA | CUDA graph | Cluster scheduling | CompilerEquity | Health insuranceMid-level Full TimeSan Francisco, Singapore, Amsterdam1mo ago
-
Deep learning | Distributed Training | Flash Attention | Inference Optimization | Kernel FusionHybrid workSenior-level Full TimeToronto, Ontario, Canada1mo ago
-
Research Intern, Inference (Fall 2026) USD 116K-126KCUDA | Deep learning | Distributed Systems | JAX | Machine LearningHousing stipend | Open source contribution opportunitiesEntry-level InternshipSan Francisco1mo ago
-
Staff Software Engineer, Inference PLN 369K-542KAutoscaling | BF16 | Benchmarking | C++ | CUDACritical illness cover | Employee assistance programme | Family dental insurance | Family medical insurance | Generous pension contributionSenior-level Full TimeWarsaw, Poland1mo ago
-
AI Engineer USD 100K-135KAWQ | AWS | AWS EC2 | Agent Frameworks | CI/CD401k match | Health insurance | Learning and development stipend | Paid parental leave | Paid time offMid-level Full TimeRemote USA - In Tandem R1mo ago
-
Attention Mechanisms | C++ | Decoder Only | Decoder-only Transformer | GPU parallelismComprehensive benefitsSenior-level Full TimeNew York, New York, United States …1mo ago
-
Senior Quantization Engineer - Edge AI Model Optimization INR 3000K-5000KC++ | CNN | Deep learning | Embedded Systems | Generative AISenior-level Full TimeHyderabad, India1mo ago