AI Researcher - Efficient AI (Contractor)
Tasks
- Create experimental evaluation pipelines
- Develop KV cache compression approaches
- Evaluate compression and pruning methods
- Implement backpropagation free optimization methods
- Implement gradient free model merging methods
- Optimize efficient LLM inference performance
- Prototype AI model compression approaches
- Prototype speculative decoding for low latency generation
- Publish technical reports and IP disclosures
- Research model efficiency methods
Perks/Benefits
- 401k matching
- Fitness Goal Incentives
- Health, dental, vision insurance
- Hybrid work
- Life and disability insurance
- Mental health resources
- Paid parental leave
- Paid time off
Skills/Tech-stack
Agent systems | Cache Compression | DPO | Distillation | Hybrid Attention | Inference Optimization | Instruction Tuning | KV Cache Compression | KV cache | Kernel optimization | Llama.cpp | LoRA | Long Context | Long context inference | Looped Transformers | Low Latency | Low Rank Approximation | MOE | Machine Learning | Mixture of Experts | Model Compression | Model Merging | Model Pruning | Multi-Agent | Multi-Agent Systems | Post-training | PyTorch | Python | Quantization | Quantization aware training | Reinforcement Learning | SGLang | Speculative decoding | State Space Models | State-Space | TensorRT-LLM | VLLM
Education
Roles
Regions
Countries
States
Related jobs
-
Senior AI Solutions Architect - Semiconductors USD 152K-287KC plus plus | CUDA | Computational lithography | Containers | Data PipelinesBenefits | Equity | Hackathons | Technical demonstrations | TrainingSenior-level Full TimeUS, CA, Santa Clara R9d ago
-
C++ | Containerization | Deep learning | Deep learning inference | GPU OrchestrationSenior-level Full TimeUS, CA, Santa Clara R9d ago
-
Applied Agentic AI Lead, Partner Co-Design USD 224K-356KAgent systems | Compliance governance | Deep learning | Docker | Efficient Fine TuningEquity | Health benefits | Professional developmentSenior-level Full TimeUS, CA, Santa Clara R11d ago