Ryz Labs is seeking a Head of Sales to own the GTM strategy, execution, and revenue growth for our RLHF vertical. Based in San Francisco, this leader will be responsible for building and managing the ...
Quick apply
Ryz Labs is seeking a Head of Sales to own the GTM strategy, execution, and revenue growth for our RLHF vertical. Based in San Francisco, this leader will be responsible for building and managing the ...
Quick apply
Ryz Labs is seeking a Head of Sales to own the GTM strategy, execution, and revenue growth for our RLHF vertical. Based in San Francisco, this leader will be responsible for building and managing the ...
Ryz Labs is seeking a Head of Sales to own the GTM strategy, execution, and revenue growth for our RLHF vertical. Based in San Francisco, this leader will be responsible for building and managing the ...
Ryz Labs is seeking a Head of Sales to own the GTM strategy, execution, and revenue growth for our RLHF vertical. Based in San Francisco, this leader will be responsible for building and managing the ...
San Francisco, CA · On-site +1
Ryz Labs is seeking a Head of Sales to own the GTM strategy, execution, and revenue growth for our RLHF vertical. Based in San Francisco, this leader will be responsible for building and managing the ...
San Francisco, CA · On-site +1
Ryz Labs is seeking a Head of Sales to own the GTM strategy, execution, and revenue growth for our RLHF vertical. Based in San Francisco, this leader will be responsible for building and managing the ...
They combine platforms, tools and a large expert community to deliver training data, evaluation, RLHF and multilingual AI solutions for complex, high impact use cases. The work is global, fast moving ...
They combine platforms, tools and a large expert community to deliver training data, evaluation, RLHF and multilingual AI solutions for complex, high impact use cases. The work is global, fast moving ...
San Francisco, CA · On-site
... PT / SFT / RLHF / Eval 各阶段对数据的不同要求,推动数据标准持续迭代. 任职要求 • 3 年以上相关经验,做过视频 / 图像 / 多模态数据的标注或评测 ...
New
San Francisco, CA · On-site
... PT / SFT / RLHF / Eval 各阶段对数据的不同要求,推动数据标准持续迭代. 任职要求 • 3 年以上相关经验,做过视频 / 图像 / 多模态数据的标注或评测 ...
New
San Jose, CA · On-site
$250K - $350K/yr
You will build and improve RLHF pipelines, develop reward models, and design post-training workflows using approaches such as PPO, DPO, GRPO, KTO, and similar methods. You'll work closely with ...
San Jose, CA · On-site
$250K - $350K/yr
You will build and improve RLHF pipelines, develop reward models, and design post-training workflows using approaches such as PPO, DPO, GRPO, KTO, and similar methods. You'll work closely with ...
$180K - $600K/yr
You will work on the most critical post-training and reinforcement learning challenges at any given time - including reward modeling, preference optimization (RLHF/DPO), and RL for improving ...
$180K - $600K/yr
You will work on the most critical post-training and reinforcement learning challenges at any given time - including reward modeling, preference optimization (RLHF/DPO), and RL for improving ...
Durham, NC · Remote
You will play a critical role in shaping the future of AI-driven contact center platforms, combining Generative AI, GraphRAG, RLHF, and multi-agent systems to deliver highly personalized, context ...
Quick apply
Durham, NC · Remote
You will play a critical role in shaping the future of AI-driven contact center platforms, combining Generative AI, GraphRAG, RLHF, and multi-agent systems to deliver highly personalized, context ...
Preferred : • Experience with implementing LLM finetuning algorithms (such as RLHF) and modifying systems based on model architectures. • Worked on the post-training or infra team at frontier ...
Preferred : • Experience with implementing LLM finetuning algorithms (such as RLHF) and modifying systems based on model architectures. • Worked on the post-training or infra team at frontier ...
Preferred : • Demonstrated expertise in post-training methods and/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and ...
Preferred : • Demonstrated expertise in post-training methods and/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and ...
Austin, TX · On-site
$103K - $142K/yr
... and RLHF. Responsibilities : • Lead the fine-tuning process for large pre-trained models, focusing on making models behave appropriately in different contexts (e.g., following instructions ...
Austin, TX · On-site
$103K - $142K/yr
... and RLHF. Responsibilities : • Lead the fine-tuning process for large pre-trained models, focusing on making models behave appropriately in different contexts (e.g., following instructions ...
Proficiency in modern RLHF algorithms: PPO, DPO, GRPO, etc. * Hands-on experience training reward models and finetuning LLM/VLM/VLA * Knowledge of distributed RL training at scale * Proficiency with ...
Proficiency in modern RLHF algorithms: PPO, DPO, GRPO, etc. * Hands-on experience training reward models and finetuning LLM/VLM/VLA * Knowledge of distributed RL training at scale * Proficiency with ...
Responsibilities : • You will work on the most critical post-training and reinforcement learning challenges at any given time -- including reward modeling, preference optimization (RLHF/DPO), and ...
Responsibilities : • You will work on the most critical post-training and reinforcement learning challenges at any given time -- including reward modeling, preference optimization (RLHF/DPO), and ...
Collaborate closely with ML researchers to implement stable and fast versions of new finetuning recipes (like in RLHF/SFT) on different model architectures. What We'd Like to See Qualifications ...
Collaborate closely with ML researchers to implement stable and fast versions of new finetuning recipes (like in RLHF/SFT) on different model architectures. What We'd Like to See Qualifications ...
Proficiency in modern RLHF algorithms: PPO, DPO, GRPO, etc. * Hands-on experience training reward models and finetuning LLM/VLM/VLA * Knowledge of distributed RL training at scale * Proficiency with ...
Proficiency in modern RLHF algorithms: PPO, DPO, GRPO, etc. * Hands-on experience training reward models and finetuning LLM/VLM/VLA * Knowledge of distributed RL training at scale * Proficiency with ...
Proficiency in modern RLHF algorithms: PPO, DPO, GRPO, etc. * Hands-on experience training reward models and finetuning LLM/VLM/VLA * Knowledge of distributed RL training at scale * Proficiency with ...
Quick apply
Proficiency in modern RLHF algorithms: PPO, DPO, GRPO, etc. * Hands-on experience training reward models and finetuning LLM/VLM/VLA * Knowledge of distributed RL training at scale * Proficiency with ...
Required : • Hands-on experience with data generation and evaluation for LLM post-training • Experience training or fine-tuning models using SFT, instruction tuning, RLHF, DPO, or similar ...
Required : • Hands-on experience with data generation and evaluation for LLM post-training • Experience training or fine-tuning models using SFT, instruction tuning, RLHF, DPO, or similar ...
Preferred : • Demonstrated expertise in post-training methods and/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and ...
Preferred : • Demonstrated expertise in post-training methods and/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and ...
Redmond, WA · On-site
$220K - $331K/yr
You'll be embedded directly with applied researchers and ML engineers running RLHF, fine-tuning, and preference-tuning pipelines, reading model outputs, calling out what's wrong, and turning that ...
Redmond, WA · On-site
$220K - $331K/yr
You'll be embedded directly with applied researchers and ML engineers running RLHF, fine-tuning, and preference-tuning pipelines, reading model outputs, calling out what's wrong, and turning that ...
Santa Clara, CA · On-site
$122K - $168K/yr
Preferred : • Experience with RLHF or preference learning. • Experience with LLM agents or tool-using AI systems. • Multi-agent systems or long-horizon planning. • Simulation environments for ...
Santa Clara, CA · On-site
$122K - $168K/yr
Preferred : • Experience with RLHF or preference learning. • Experience with LLM agents or tool-using AI systems. • Multi-agent systems or long-horizon planning. • Simulation environments for ...
| Aspect | Rlhf | Rn |
|---|---|---|
| Required Credentials | Licensed healthcare professional, often with specialized training in mental health or behavioral health | Licensed practical nurse or registered nurse, with nursing licensure and possibly additional certifications |
| Work Environment | Behavioral health facilities, clinics, hospitals, or community health settings | Hospitals, clinics, long-term care facilities, and community health settings |
| Employer & Industry Usage | Behavioral health and mental health services | General healthcare and nursing services |
| Common Search & Comparison | Rlhf vs Rn | Rlhf vs Rn |
While Rlhf (Registered Licensed Mental Health Facilitator) focuses on mental health support and behavioral health interventions, Rn (Registered Nurse) provides broader nursing care across various medical settings. Both roles require licensure, but Rlhf specializes in mental health, whereas Rn covers general patient care.
An RLHF (Reinforcement Learning with Human Feedback) job involves training AI models using human feedback to improve their responses. Professionals in this role analyze model outputs, provide evaluations, and refine AI behavior through reinforcement learning techniques. These roles are common in AI research, content moderation, and chatbot development.

Sourced by ZipRecruiter
Software development
1 - 10 Employees
Los Angeles, CA, US
2021