... online RL for large multimodal policies • Close the sim-2-real gap through tactile sensing ... models with RLHF, DPO, GRPO, or related methods • Familiarity with tactile sensing, multi ...
... online RL for large multimodal policies • Close the sim-2-real gap through tactile sensing ... models with RLHF, DPO, GRPO, or related methods • Familiarity with tactile sensing, multi ...
Research Scientist, RL for Dexterous Manipulation, Atlas
Waltham, MA · On-site +1
$175K - $220K/yr
Research reward modeling, and offline-to-online RL for large multimodal policies * Close the sim-2 ... Experience fine-tuning foundation models with RLHF, DPO, GRPO, or related methods * Familiarity ...
Research Scientist, RL for Dexterous Manipulation, Atlas
Waltham, MA · On-site +1
$175K - $220K/yr
Research reward modeling, and offline-to-online RL for large multimodal policies * Close the sim-2 ... Experience fine-tuning foundation models with RLHF, DPO, GRPO, or related methods * Familiarity ...
Research Scientist, RL for Dexterous Manipulation, Atlas
Waltham, MA · On-site
$175K - $220K/yr
Research reward modeling, and offline-to-online RL for large multimodal policies * Close the sim-2 ... Experience fine-tuning foundation models with RLHF, DPO, GRPO, or related methods * Familiarity ...
Research Scientist, RL for Dexterous Manipulation, Atlas
Waltham, MA · On-site
$175K - $220K/yr
Research reward modeling, and offline-to-online RL for large multimodal policies * Close the sim-2 ... Experience fine-tuning foundation models with RLHF, DPO, GRPO, or related methods * Familiarity ...
Online Rlhf information
See Boston, MA salary details
$19K - $25.8K
12% of jobs
$30.7K is the 25th percentile. Wages below this are outliers.
$25.8K - $32.5K
18% of jobs
$32.5K - $39.3K
15% of jobs
The median wage is $40.3K / yr.
$39.3K - $46.1K
33% of jobs
$46.1K - $52.8K
12% of jobs
$52.8K - $59.6K
0% of jobs
$59.6K - $66.4K
0% of jobs
$66.4K - $73.1K
0% of jobs
$73.1K - $79.9K
0% of jobs
$79.9K - $86.7K
0% of jobs
$86.7K - $93.4K
9% of jobs
$19K
$44.1K
$93.4K
How much do online rlhf jobs pay per year?
What are some common challenges faced by Online RLHF (Reinforcement Learning from Human Feedback) specialists when collaborating with cross-functional teams?
What is the difference between Online Rlhf vs Online Rlhf?
| Aspect | Online Rlhf | Online Rlhf |
|---|---|---|
| Credentials | Typically requires certification in online health coaching or related fields | Typically requires certification in online health coaching or related fields |
| Work Environment | Remote, online platform-based | Remote, online platform-based |
| Industry Usage | Common in health and wellness sectors | Common in health and wellness sectors |
| Job Focus | Providing health guidance and support online | Providing health guidance and support online |
Online Rlhf and Online Rlhf are the same role, often used interchangeably. Both involve providing health and wellness support remotely, requiring similar certifications and working within the online health industry. The key difference is often in terminology rather than job function.
What are Online RLHF jobs?
What are the key skills and qualifications needed to thrive as an Online RLHF (Reinforcement Learning from Human Feedback) Specialist, and why are they important?
Full-time
Re-posted 21 days ago
Job description
Boston Dynamics is a robotics company known for its innovative humanoid robots. They are seeking a Research Scientist to lead projects on reinforcement learning for dexterous manipulation tasks, focusing on visual sim-to-real transfer and post-training of Vision-Language-Action models.
Responsibilities:
• Develop novel algorithms for visual sim-to-real transfer with photorealistic rendering
• Design post-training recipes that improve pretrained VLA models on manipulation tasks
• Research reward modeling, and offline-to-online RL for large multimodal policies
• Close the sim-2-real gap through tactile sensing, vision, and system identification
• Train policies that generalize across objects, scenes, and embodiments
Qualifications:
Required:
• PhD, in ML, Robotics, or a related field or a MS with 3+ years of experience
• Track record of first-author publications at top venues (CoRL, RSS, ICLR, NeuRIPS)
• Demonstrated experience training policies for dexterous or contact-rich manipulation
• Hands-on experience with VLA models, diffusion policies, or large behavior models
• Proficient in PyTorch and/or JAX, with experience training models at scale
• Strong software fundamentals and the ability to ship research code that runs reliably
Preferred:
• Deployed vision-based manipulation policies on physical robots
• Deep knowledge of sim-to-real transfer techniques and photorealistic rendering
• Built training pipelines that combine RL, imitation learning, and large-scale pretraining
• Experience fine-tuning foundation models with RLHF, DPO, GRPO, or related methods
• Familiarity with tactile sensing, multi-fingered hands, or bimanual coordination
Company:
Boston Dynamics is an engineering company that specializes in building dynamic robots and software for human simulation. It is a sub-organization of Hyundai Motor Company. Founded in 1992, the company is headquartered in Waltham, USA, with a team of 501-1000 employees. The company is currently Late Stage.
About Boston Dynamics
Sourced by ZipRecruiter
Industry
Industrial automation equipment manufacturing
Company size
51 - 200 Employees
Headquarters location
Waltham, MA, US
Year founded
1992