1

Online Rlhf Jobs (NOW HIRING)

Apply Reinforcement Learning (RLVR, RLHF), Direct Preference Optimization (DPO), and customer ... well as online experiments. About the team Core Search builds the next-generation LLM-powered ...

Apply Reinforcement Learning (RLVR, RLHF), Direct Preference Optimization (DPO), and customer ... well as online experiments. About the team Core Search builds the next-generation LLM-powered ...

... online experimentation. Personalization is a first-class objective on this team. You will build ... GRPO, DPO, RLHF), knowledge distillation, and listwise ranking losses (e.g. LambdaLoss, ListNet ...

Senior Inference Engineer, AGI

Sunnyvale, CA · On-site

$122K - $168K/yr

... RL/RLHF/RLAIF) Ensure train/serve consistency - that the inference path used in RL and evaluation faithfully matches production online behavior (e.g., parity across sampling and logit processing ...

Senior Inference Engineer, AGI

Sunnyvale, CA · On-site

$122K - $168K/yr

... RL/RLHF/RLAIF) Ensure train/serve consistency - that the inference path used in RL and evaluation faithfully matches production online behavior (e.g., parity across sampling and logit processing ...

Define experimentation strategies, offline evaluation frameworks, and online/offline metric ... or RLHF. * Strong engineering skills in Python and Scala, with experience building large-scale ...

Define experimentation strategies, offline evaluation frameworks, and online/offline metric ... or RLHF. * Strong engineering skills in Python and Scala, with experience building large-scale ...

Senior Machine Learning Engineer, Payments

$107K - $146K/yr

Proven mastery of modern AI/LLM workflows - prompt engineering, fine tuning (LoRA, RLHF), hallucination mitigation, safety guardrails, and rigorous online/offline testing to minimize training ...

Senior Inference Engineer, AGI

Sunnyvale, CA

$122K - $168K/yr

... RL/RLHF/RLAIF) Ensure train/serve consistency - that the inference path used in RL and evaluation faithfully matches production online behavior (e.g., parity across sampling and logit processing ...

Showing results 21-40

Online Rlhf information

See salary details

$17.5K

$40.6K

$86K

How much do online rlhf jobs pay per year?

As of Sep 13, 2026, the average yearly pay for online rlhf in the United States is $40,596.00, according to ZipRecruiter salary data. Most workers in this role earn between $25,000.00 and $43,500.00 per year, depending on experience, location, and employer.

What is an online RLHF?

Online RLHF (Reinforcement Learning from Human Feedback) jobs typically involve helping to train AI models by providing human feedback on their outputs. Workers in these roles might review model responses, rate the quality of generated text, or suggest improvements to help the AI learn to produce better results. These jobs are often remote and can be done part-time or as contract work. They play a crucial role in improving the safety, usefulness, and accuracy of AI systems by aligning them more closely with human preferences.

What are some common challenges faced by online RLHF specialists when collaborating with cross-functional teams?

Online RLHF specialists often work closely with machine learning engineers, data annotators, and product managers. A common challenge is ensuring that feedback from human annotators is accurately integrated into model training, which requires clear communication and well-defined annotation guidelines. Additionally, balancing the pace of model updates with the need for high-quality human feedback can be demanding. Effective collaboration and regular syncs are essential to maintain alignment and achieve project goals.

What are the key skills and qualifications needed to thrive as an online RLHF specialist, and why are they important?

To thrive as an Online RLHF Specialist, you need a strong background in machine learning, reinforcement learning, and data analysis, typically supported by a degree in computer science or a related field. Familiarity with technical tools like Python, PyTorch or TensorFlow, and experience with human feedback systems or annotation platforms are highly valuable. Strong problem-solving, attention to detail, and the ability to communicate complex concepts clearly are crucial soft skills. These qualifications ensure the effective training and evaluation of AI models, leading to more accurate and reliable machine learning systems.

What is the difference between Online Rlhf vs Online Rlhf?

AspectOnline RlhfOnline Rlhf
CredentialsTypically requires certification in online health coaching or related fieldsTypically requires certification in online health coaching or related fields
Work EnvironmentRemote, online platform-basedRemote, online platform-based
Industry UsageCommon in health and wellness sectorsCommon in health and wellness sectors
Job FocusProviding health guidance and support onlineProviding health guidance and support online

Online Rlhf and Online Rlhf are the same role, often used interchangeably. Both involve providing health and wellness support remotely, requiring similar certifications and working within the online health industry. The key difference is often in terminology rather than job function.

More about Online Rlhf jobs

What cities are hiring for Online Rlhf jobs?

Cities with the most Online Rlhf job openings:

What are the most commonly searched types of Rlhf jobs?

The most popular types of Rlhf jobs are:

What states have the most Online Rlhf jobs?

States with the most job openings for Online Rlhf jobs include:

Infographic showing various Online Rlhf job openings in the United States as of September 2026, with employment types broken down into 1% As Needed, 61% Full Time, 35% Part Time, 1% Temporary, and 2% Contract. Highlights an 80% Physical, 1% Hybrid, and 19% Remote job distribution, with an average salary of $40,596 per year, or $19.5 per hour.

Research Scientist, RL for Dexterous Manipulation, Atlas

Waltham, MA • On-site

$175K - $220K/yr

Full-time

Medical, Dental, Vision, Retirement, PTO

Re-posted 4 days ago


Job description

Are you passionate about using reinforcement learning to solve dexterous manipulation tasks? As an RL Research Scientist, you'll lead research projects on visual sim-to-real transfer and post-training of Vision-Language-Action (VLA) models. Your job is to turn unlabeled data from simulation or the real-world into robust real-world manipulation skills. You should push the frontier of what bimanual and multi-fingered systems can do in unstructured environments.

In this role, you will:

  • Develop novel algorithms for visual sim-to-real transfer with photorealistic rendering

  • Design post-training recipes that improve pretrained VLA models on manipulation tasks

  • Research reward modeling, and offline-to-online RL for large multimodal policies

  • Close the sim-2-real gap through tactile sensing, vision, and system identification

  • Train policies that generalize across objects, scenes, and embodiments

Required Qualifications:

  • PhD, in ML, Robotics, or a related field or a MS with 3+ years of experience

  • Track record of first-author publications at top venues (CoRL, RSS, ICLR, NeuRIPS)

  • Demonstrated experience training policies for dexterous or contact-rich manipulation

  • Hands-on experience with VLA models, diffusion policies, or large behavior models

  • Proficient in PyTorch and/or JAX, with experience training models at scale

  • Strong software fundamentals and the ability to ship research code that runs reliably

The ideal candidate has:

  • Deployed vision-based manipulation policies on physical robots

  • Deep knowledge of sim-to-real transfer techniques and photorealistic rendering

  • Built training pipelines that combine RL, imitation learning, and large-scale pretraining

  • Experience fine-tuning foundation models with RLHF, DPO, GRPO, or related methods

  • Familiarity with tactile sensing, multi-fingered hands, or bimanual coordination

Why join us?

  • Direct impact on the policies powering our humanoid robots in the real world

  • Access to a world-class fleet of robots, simulation infrastructure, and compute

  • The chance to define what's possible for dexterous manipulation at scale

The pay range for this position is between $175,000 to $220,000 annually. Base pay will depend on multiple individualized factors including, but not limited to internal equity, job related knowledge, skills and experience. This range represents a good faith estimate of compensation at the time of posting. Boston Dynamics offers a generous Benefits package including medical, dental vision, 401(k), paid time off and a annual bonus structure. Additional details regarding these benefit plans will be provided if an employee receives an offer for employment.

#LI-JM1