Applied Research Engineer Location - San Francisco, CA - Hybrid Job Type: Full Time Role Overview ... Drive advances in AI alignment , developing state-of-the-art techniques like RLHF and new methods ...
Quick apply
Applied Research Engineer Location - San Francisco, CA - Hybrid Job Type: Full Time Role Overview ... Drive advances in AI alignment , developing state-of-the-art techniques like RLHF and new methods ...
Quick apply
Applied Research Engineer Location - San Francisco, CA - Hybrid Job Type: Full Time Role Overview ... Drive advances in AI alignment , developing state-of-the-art techniques like RLHF and new methods ...
Santa Clara, CA · On-site
$122K - $168K/yr
Learning from human or AI feedback (RLHF / RLAIF). * Agent training pipelines built on top of our ... The base salary range for this full-time position is $244,140 - $413,160, in addition to bonus ...
Santa Clara, CA · On-site
$122K - $168K/yr
Learning from human or AI feedback (RLHF / RLAIF). * Agent training pipelines built on top of our ... The base salary range for this full-time position is $244,140 - $413,160, in addition to bonus ...
Santa Clara, CA · On-site
$122K - $168K/yr
Learning from human or AI feedback (RLHF / RLAIF). * Agent training pipelines built on top of our ... The base salary range for this full-time position is $244,140 - $413,160, in addition to bonus ...
Santa Clara, CA · On-site
$122K - $168K/yr
Learning from human or AI feedback (RLHF / RLAIF). * Agent training pipelines built on top of our ... The base salary range for this full-time position is $244,140 - $413,160, in addition to bonus ...
Los Gatos, CA · On-site
$238K/yr
Experience with quality and evaluation metrics for complex ML products, RLHF, human-in-the-loop ... Full-time hourly employees accrue 35 days annually for paid time off to be used for vacation ...
Los Gatos, CA · On-site
$238K/yr
Experience with quality and evaluation metrics for complex ML products, RLHF, human-in-the-loop ... Full-time hourly employees accrue 35 days annually for paid time off to be used for vacation ...
Experience with quality and evaluation metrics for complex ML products, RLHF, human-in-the-loop ... Full-time hourly employees accrue 35 days annually for paid time off to be used for vacation ...
Experience with quality and evaluation metrics for complex ML products, RLHF, human-in-the-loop ... Full-time hourly employees accrue 35 days annually for paid time off to be used for vacation ...
... RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation ... Ideally you'd have: * 3+ years of full-time engineering experience, post-graduation with ...
... RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation ... Ideally you'd have: * 3+ years of full-time engineering experience, post-graduation with ...
San Francisco, CA · On-site
San Francisco Bay Area Type: Full-Time Compensation: Competitive salary + meaningful equity ... RLHF, DPO, PPO) * Establish scalable evaluation harnesses for LLM and agent performance, including ...
San Francisco, CA · On-site
San Francisco Bay Area Type: Full-Time Compensation: Competitive salary + meaningful equity ... RLHF, DPO, PPO) * Establish scalable evaluation harnesses for LLM and agent performance, including ...
San Francisco Bay Area Type: Full-Time Compensation: Competitive salary + meaningful equity ... RLHF, DPO, PPO) * Establish scalable evaluation harnesses for LLM and agent performance, including ...
Quick apply
San Francisco Bay Area Type: Full-Time Compensation: Competitive salary + meaningful equity ... RLHF, DPO, PPO) * Establish scalable evaluation harnesses for LLM and agent performance, including ...
San Francisco, CA · On-site
$195K - $365K/yr
Applying Reinforcement Learning (RL) techniques (like RLHF or GRPO) to improve conversational ... Retirement Planning: 401(k) plan for full-time employees with company matching. * Paid Time Off:
San Francisco, CA · On-site
$195K - $365K/yr
Applying Reinforcement Learning (RL) techniques (like RLHF or GRPO) to improve conversational ... Retirement Planning: 401(k) plan for full-time employees with company matching. * Paid Time Off:
Mountain View, CA · On-site
$35 - $50/hr
You will be paired with a full-time employee as your mentor, working together to explore, from zero ... RLHF, DPO, PPO, GRPO, etc.) * Excellent taste in model behavior: able to reason about what "good ...
Mountain View, CA · On-site
$35 - $50/hr
You will be paired with a full-time employee as your mentor, working together to explore, from zero ... RLHF, DPO, PPO, GRPO, etc.) * Excellent taste in model behavior: able to reason about what "good ...
Mountain View, CA · On-site
$150K - $230K/yr
... RLHF, PPO, GRPO, DPO, and related methods). * Design, build, and curate the data that drives each ... Team activity budget The US base salary range for this full-time position is listed below. Pay may ...
Mountain View, CA · On-site
$150K - $230K/yr
... RLHF, PPO, GRPO, DPO, and related methods). * Design, build, and curate the data that drives each ... Team activity budget The US base salary range for this full-time position is listed below. Pay may ...
Mountain View, CA · On-site
$150K - $230K/yr
... RLHF, PPO, GRPO, DPO, and related methods). * Design, build, and curate the data that drives each ... Team activity budget The US base salary range for this full-time position is listed below. Pay may ...
Quick apply
Mountain View, CA · On-site
$150K - $230K/yr
... RLHF, PPO, GRPO, DPO, and related methods). * Design, build, and curate the data that drives each ... Team activity budget The US base salary range for this full-time position is listed below. Pay may ...
San Jose, CA · On-site
$180K - $450K/yr
RL, RLHF/RLAIF, reward modeling and graders, synthetic data pipelines, or building agentic/tool ... Compensation The US base salary range for this full-time position is between $180,000 - $450,000 ...
San Jose, CA · On-site
$180K - $450K/yr
RL, RLHF/RLAIF, reward modeling and graders, synthetic data pipelines, or building agentic/tool ... Compensation The US base salary range for this full-time position is between $180,000 - $450,000 ...
San Jose, CA · On-site
$180K - $450K/yr
RL, RLHF/RLAIF, reward modeling and graders, synthetic data pipelines, or building agentic/tool ... Compensation The US base salary range for this full-time position is between $180,000 - $450,000 ...
San Jose, CA · On-site
$180K - $450K/yr
RL, RLHF/RLAIF, reward modeling and graders, synthetic data pipelines, or building agentic/tool ... Compensation The US base salary range for this full-time position is between $180,000 - $450,000 ...
San Jose, CA · On-site
$200K - $400K/yr
Experience with post-training techniques such as RLHF, reward modeling, or alignment methods ... full-time position is between $200,000 - $400,000 The pay offered for this position may vary based ...
San Jose, CA · On-site
$200K - $400K/yr
Experience with post-training techniques such as RLHF, reward modeling, or alignment methods ... full-time position is between $200,000 - $400,000 The pay offered for this position may vary based ...
San Jose, CA · On-site
$200K - $400K/yr
Experience with post-training techniques such as RLHF, reward modeling, or alignment methods ... full-time position is between $200,000 - $400,000 The pay offered for this position may vary based ...
San Jose, CA · On-site
$200K - $400K/yr
Experience with post-training techniques such as RLHF, reward modeling, or alignment methods ... full-time position is between $200,000 - $400,000 The pay offered for this position may vary based ...
San Jose, CA · On-site
$200K - $400K/yr
Experience with post-training techniques such as RLHF, reward modeling, or alignment methods ... full-time position is between $200,000 - $400,000 The pay offered for this position may vary based ...
San Jose, CA · On-site
$200K - $400K/yr
Experience with post-training techniques such as RLHF, reward modeling, or alignment methods ... full-time position is between $200,000 - $400,000 The pay offered for this position may vary based ...
San Jose, CA · On-site
$200K - $400K/yr
Experience with post-training techniques such as RLHF, reward modeling, or alignment methods ... full-time position is between $200,000 - $400,000 The pay offered for this position may vary based ...
San Jose, CA · On-site
$200K - $400K/yr
Experience with post-training techniques such as RLHF, reward modeling, or alignment methods ... full-time position is between $200,000 - $400,000 The pay offered for this position may vary based ...
Mountain View, CA · On-site
$35 - $50/hr
You will be paired with a full-time employee as your mentor, working together to explore, from zero ... RLHF, DPO, PPO, GRPO, etc.) * Excellent taste in model behavior: able to reason about what "good ...
Quick apply
Mountain View, CA · On-site
$35 - $50/hr
You will be paired with a full-time employee as your mentor, working together to explore, from zero ... RLHF, DPO, PPO, GRPO, etc.) * Excellent taste in model behavior: able to reason about what "good ...
Mountain View, CA · On-site
$35 - $50/hr
You will be paired with a full-time employee as your mentor, working together to explore, from zero ... RLHF, DPO, PPO, GRPO, etc.) * Excellent taste in model behavior: able to reason about what "good ...
Mountain View, CA · On-site
$35 - $50/hr
You will be paired with a full-time employee as your mentor, working together to explore, from zero ... RLHF, DPO, PPO, GRPO, etc.) * Excellent taste in model behavior: able to reason about what "good ...
$12.98 - $14.07
11% of jobs
$15.06 is the 25th percentile. Wages below this are outliers.
$14.07 - $15.17
16% of jobs
The median wage is $16.26 / hr.
$15.17 - $16.26
23% of jobs
$16.26 - $17.35
11% of jobs
$17.35 - $18.44
11% of jobs
$18.64 is the 75th percentile. Wages above this are outliers.
$18.44 - $19.54
21% of jobs
$19.54 - $20.63
3% of jobs
$20.63 - $21.72
1% of jobs
$21.72 - $22.81
1% of jobs
$22.81 - $23.91
1% of jobs
$23.91 - $25
1% of jobs
$12
$17
$25
The most popular types of Rlhf jobs are:
The top searched job categories for Full Time Rlhf jobs are:

Full-time
Re-posted 23 days ago
Position: Applied Research Engineer
Location - San Francisco, CA - Hybrid
Job Type: Full Time
Job Description:
Role OverviewAs an Applied Research Engineer at Labelbox, you’ll play a critical role in shaping the future of human-in-the-loop AI systems. You’ll design and implement advanced methods to align human feedback with the training of cutting-edge AI models, including techniques like Reinforcement Learning from Human Feedback (RLHF), Direct Preference Optimization (DPO), and other alignment strategies. You’ll also develop innovative tools to measure and enhance the quality of human-generated data and build AI-assisted systems that streamline and improve the data labeling process.
Your work will directly influence the performance and reliability of frontier models by ensuring they reflect human preferences more accurately. By bridging research with real-world applications, you’ll help bring scalable, impactful alignment solutions into production for some of the world’s most advanced AI developers.
Your Impact