Senior Deep Learning Scientist, Multimodal Agentic RL
Tasks
- Advance post training and alignment methods
- Collect develop and benchmark multimodal datasets
- Develop train fine tune deploy large language models for agentic systems
- Evaluate model accuracy safety and task success
- Implement reinforcement learning algorithms and reward design
- Manage dataset versioning experiment tracking and evaluation pipelines
- Plan tool execution and long horizon task completion
- Research agentic reasoning and grounded perception
Perks/Benefits
Skills/Tech-stack
Dataset versioning | Deep learning | Distributed Training | Experiment tracking | Instruction Tuning | MDP | Mixture of Experts | Policy Optimization | Preference optimization | PyTorch | Python | RLHF | Reinforcement Learning | Reward Design | Transformers
Education
Regions
Countries
States
Related jobs
- No jobs found.