... RLHF/RLAIF, online RL, and model-based data improvement. • Design the systems abstractions that connect research ideas to production-scale RL runs: trainers, rollout workers, reward models ...
... RLHF/RLAIF, online RL, and model-based data improvement. • Design the systems abstractions that connect research ideas to production-scale RL runs: trainers, rollout workers, reward models ...
Member of Technical Staff - RL Research (Experienced)
$300K - $500K/yr
Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement. * Design the systems abstractions that connect research ...
Member of Technical Staff - RL Research (Experienced)
$300K - $500K/yr
Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement. * Design the systems abstractions that connect research ...
Experience with agent evaluation, offline/online experiments, and human feedback loops in production. * Direct experience with RLHF, RLAIF, DPO, PPO, GRPO, or related optimization techniques. * Prior ...
Experience with agent evaluation, offline/online experiments, and human feedback loops in production. * Direct experience with RLHF, RLAIF, DPO, PPO, GRPO, or related optimization techniques. * Prior ...
... RLHF/RLAIF, online RL, and model-based data improvement. • Design the systems abstractions that connect research ideas to production-scale RL runs: trainers, rollout workers, reward models ...
... RLHF/RLAIF, online RL, and model-based data improvement. • Design the systems abstractions that connect research ideas to production-scale RL runs: trainers, rollout workers, reward models ...
Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement. * Design the systems abstractions that connect research ...
Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement. * Design the systems abstractions that connect research ...
Experience with agent evaluation, offline/online experiments, and human feedback loops in production. * Direct experience with RLHF, RLAIF, DPO, PPO, GRPO, or related optimization techniques. * Prior ...
Experience with agent evaluation, offline/online experiments, and human feedback loops in production. * Direct experience with RLHF, RLAIF, DPO, PPO, GRPO, or related optimization techniques. * Prior ...
Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants. * Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF. * Set up ...
Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants. * Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF. * Set up ...
Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants. * Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF. * Set up ...
Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants. * Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF. * Set up ...
Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants. * Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF. * Set up ...
Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants. * Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF. * Set up ...
Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants. * Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF. * Set up ...
Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants. * Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF. * Set up ...
Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement. * Design the systems abstractions that connect research ...
Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement. * Design the systems abstractions that connect research ...
Apply Reinforcement Learning (RLVR, RLHF), Direct Preference Optimization (DPO), and customer ... well as online experiments. About the team Core Search builds the next-generation LLM-powered ...
Apply Reinforcement Learning (RLVR, RLHF), Direct Preference Optimization (DPO), and customer ... well as online experiments. About the team Core Search builds the next-generation LLM-powered ...
Apply Reinforcement Learning (RLVR, RLHF), Direct Preference Optimization (DPO), and customer ... well as online experiments. About the team Core Search builds the next-generation LLM-powered ...
Apply Reinforcement Learning (RLVR, RLHF), Direct Preference Optimization (DPO), and customer ... well as online experiments. About the team Core Search builds the next-generation LLM-powered ...
Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants. * Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF. * Set up ...
Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants. * Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF. * Set up ...
Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants. * Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF. * Set up ...
Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants. * Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF. * Set up ...
Software Engineer - Global E-Commerce AI Search Infrastructure (TikTok Shop)
Seattle, WA · On-site
$148K - $300K/yr
... latency online services. - Proficiency in C++, Go, or Java (C++ preferred); strong systems ... SFT, RL (RLHF / PPO / GRPO), distillation, or reward model design. - Research publications at ...
Software Engineer - Global E-Commerce AI Search Infrastructure (TikTok Shop)
Seattle, WA · On-site
$148K - $300K/yr
... latency online services. - Proficiency in C++, Go, or Java (C++ preferred); strong systems ... SFT, RL (RLHF / PPO / GRPO), distillation, or reward model design. - Research publications at ...
Senior Applied Scientist
Seattle, WA · On-site
At Amazon Selection and Catalog Systems (ASCS), our mission is to power the online buying ... tuning, RLHF, or agentic architectures Amazon is an equal opportunity employer and does not ...
Senior Applied Scientist
Seattle, WA · On-site
At Amazon Selection and Catalog Systems (ASCS), our mission is to power the online buying ... tuning, RLHF, or agentic architectures Amazon is an equal opportunity employer and does not ...
Senior Software Engineer - Global E-Commerce AI Search Infrastructure (TikTok Shop)
Seattle, WA · On-site
$202K - $368K/yr
... latency online services. - Proficiency in C++, Go, or Java (C++ preferred); strong systems ... SFT, RL (RLHF / PPO / GRPO), distillation, or reward model design. - Research publications at ...
Senior Software Engineer - Global E-Commerce AI Search Infrastructure (TikTok Shop)
Seattle, WA · On-site
$202K - $368K/yr
... latency online services. - Proficiency in C++, Go, or Java (C++ preferred); strong systems ... SFT, RL (RLHF / PPO / GRPO), distillation, or reward model design. - Research publications at ...
Apply Reinforcement Learning (RLVR, RLHF), Direct Preference Optimization (DPO), and customer ... well as online experiments. About the team Core Search builds the next-generation LLM-powered ...
Apply Reinforcement Learning (RLVR, RLHF), Direct Preference Optimization (DPO), and customer ... well as online experiments. About the team Core Search builds the next-generation LLM-powered ...
Applied Scientist
Seattle, WA · On-site
At Amazon Selection and Catalog Systems (ASCS), our mission is to power the online buying ... tuning, RLHF, prompt engineering, or agentic architectures - Experience with LLM/VLM serving ...
Applied Scientist
Seattle, WA · On-site
At Amazon Selection and Catalog Systems (ASCS), our mission is to power the online buying ... tuning, RLHF, prompt engineering, or agentic architectures - Experience with LLM/VLM serving ...
Online Rlhf information
See Seattle, WA salary details
$19.9K - $27K
12% of jobs
$32.1K is the 25th percentile. Wages below this are outliers.
$27K - $34.1K
18% of jobs
$34.1K - $41.2K
15% of jobs
The median wage is $42.2K / yr.
$41.2K - $48.3K
33% of jobs
$48.3K - $55.3K
12% of jobs
$55.3K - $62.4K
0% of jobs
$62.4K - $69.5K
0% of jobs
$69.5K - $76.6K
0% of jobs
$76.6K - $83.7K
0% of jobs
$83.7K - $90.8K
0% of jobs
$90.8K - $97.9K
9% of jobs
$19.9K
$46.2K
$97.9K
How much do online rlhf jobs pay per year?
What are some common challenges faced by online RLHF specialists when collaborating with cross-functional teams?
What is the difference between Online Rlhf vs Online Rlhf?
| Aspect | Online Rlhf | Online Rlhf |
|---|---|---|
| Credentials | Typically requires certification in online health coaching or related fields | Typically requires certification in online health coaching or related fields |
| Work Environment | Remote, online platform-based | Remote, online platform-based |
| Industry Usage | Common in health and wellness sectors | Common in health and wellness sectors |
| Job Focus | Providing health guidance and support online | Providing health guidance and support online |
Online Rlhf and Online Rlhf are the same role, often used interchangeably. Both involve providing health and wellness support remotely, requiring similar certifications and working within the online health industry. The key difference is often in terminology rather than job function.
What is an online RLHF?
What are the key skills and qualifications needed to thrive as an online RLHF specialist, and why are they important?

Full-time
Re-posted 3 days ago
Job description
Nuance Labs is building innovative AI avatars with emotional intelligence, and they are seeking a Member of Technical Staff to lead reinforcement learning and post-training for large-scale models. The role involves developing and optimizing systems for RL methods and ensuring the scalability and reliability of training systems to enhance interactive behavior and performance.
Responsibilities:
• Build Nuance’s RL/post-training stack from 0→1: rollout generation, policy optimization, reward/reference model serving, data feedback loops, evaluation, checkpointing, observability, and debugging.
• Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement.
• Design the systems abstractions that connect research ideas to production-scale RL runs: trainers, rollout workers, reward models, evaluators, data queues, experience buffers, and checkpoint promotion.
• Build evaluation and feedback loops for omni behavior: turn-taking, interruption, timing, emotional response, audiovisual coherence, instruction following, and real-time interaction quality.
• Optimize the end-to-end post-training loop across rollout throughput, serving latency, GPU utilization, policy update efficiency, queueing, checkpoint overhead, and research iteration speed.
• Evolve the platform as algorithms, model architectures, reward definitions, data sources, and evaluation methods change.
Qualifications:
Required:
• Significant hands-on experience with RL, RLHF, RLAIF, post-training, alignment, or large-scale fine-tuning for modern foundation models.
• Deep understanding of RL/post-training methods: policy optimization, reward modeling, preference optimization, rejection sampling, KL control, evaluation, and data feedback loops.
• A track record reasoning about model behavior and training dynamics: reward hacking, unstable rewards, distribution shift, stale policies, mode collapse, over-optimization, noisy preferences, and evaluation mismatch.
• Proven experience building or operating RL/post-training pipelines at scale with frameworks such as verl, ms-swift, OpenRLHF, or equivalent internal systems, including integration with rollout serving systems such as vLLM.
• Experience with large-scale training or inference systems, including rollout generation, model serving, batching, queueing, GPU utilization, checkpointing, and debugging.
• Understanding of omni post-training for real-time audio-video-language interaction: temporal alignment, interruption, emotional response, and multimodal evaluation.
• Strong software engineering fundamentals, curiosity, and adaptability to new RL algorithms, model architectures, serving systems, evaluation methods, and research ideas.
Preferred:
• Prior 0→1 experience building post-training systems, RL pipelines, agent training systems, evaluation platforms, or large-scale model improvement loops.
• Experience with PPO, GRPO, DPO, online RL, RLHF/RLAIF, reward modeling, preference data, synthetic data generation, or model-based data improvement.
• Experience with omni or multimodal post-training for audio-video-language models, especially long-context or real-time interactive systems.
• Experience scaling mixed training/inference workloads across large GPU clusters.
• Experience with adjacent areas such as distributed pretraining, data infrastructure, inference serving, simulation, human/AI feedback collection, or evaluation infrastructure.
• Publications or substantial open-source contributions in RL, post-training, alignment, evaluation, ML systems, or model behavior.
Company:
Nuance Labs an AI research company is developing the first human foundation model that understands and displays emotion in real time. Founded in 2024, the company is headquartered in Seattle, USA, with a team of 11-50 employees. The company is currently Early Stage.