SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Machine Learning Engineer
Rockville, MD · On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Machine Learning Engineer
Rockville, MD · On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Machine Learning Engineer
Annapolis, MD · On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Machine Learning Engineer
Annapolis, MD · On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Machine Learning Engineer
Annandale, VA · On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Machine Learning Engineer
Annandale, VA · On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Machine Learning Engineer
Gaithersburg, MD · On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Machine Learning Engineer
Gaithersburg, MD · On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Machine Learning Engineer
Largo, MD · On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Machine Learning Engineer
Largo, MD · On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Machine Learning Engineer
Bowie, MD · On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Machine Learning Engineer
Bowie, MD · On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Machine Learning Engineer
Adelphi, MD · On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Machine Learning Engineer
Adelphi, MD · On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Research & Innovation Impact Stay ahead of industry advancements in generative AI, foundational model tech, self-supervised learning, optimization techniques, RLHF, and robustness/interpretability.
Research & Innovation Impact Stay ahead of industry advancements in generative AI, foundational model tech, self-supervised learning, optimization techniques, RLHF, and robustness/interpretability.
AI Development Lead
Mclean, VA · On-site
Research & Innovation Impact Stay ahead of industry advancements in generative AI, foundational model tech, self-supervised learning, optimization techniques, RLHF, and robustness/interpretability.
AI Development Lead
Mclean, VA · On-site
Research & Innovation Impact Stay ahead of industry advancements in generative AI, foundational model tech, self-supervised learning, optimization techniques, RLHF, and robustness/interpretability.
AI Development Lead
Mclean, VA · On-site
Research & Innovation Impact Stay ahead of industry advancements in generative AI, foundational model tech, self-supervised learning, optimization techniques, RLHF, and robustness/interpretability.
Quick apply
AI Development Lead
Mclean, VA · On-site
Research & Innovation Impact Stay ahead of industry advancements in generative AI, foundational model tech, self-supervised learning, optimization techniques, RLHF, and robustness/interpretability.
Software Engineer, Identity
Washington, DC · On-site
To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations.
Software Engineer, Identity
Washington, DC · On-site
To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations.
To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations.
To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations.
AI Data Science Expert - Remote
Washington, DC · Remote
$100 - $200/hr
Familiarity with prompt engineering, AI output evaluation, fact-checking, or RLHF is a plus. * Master's, MBA, PhD, or other advanced degree is preferred.
Quick apply
AI Data Science Expert - Remote
Washington, DC · Remote
$100 - $200/hr
Familiarity with prompt engineering, AI output evaluation, fact-checking, or RLHF is a plus. * Master's, MBA, PhD, or other advanced degree is preferred.
Rlhf information
What is an RLHF job?
An RLHF (Reinforcement Learning with Human Feedback) job involves training AI models using human feedback to improve their responses. Professionals in this role analyze model outputs, provide evaluations, and refine AI behavior through reinforcement learning techniques. These roles are common in AI research, content moderation, and chatbot development.
What are the key skills and qualifications needed to thrive as a Reinforcement Learning from Human Feedback (RLHF) engineer, and why are they important?
What are some common challenges faced by professionals working in Reinforcement Learning from Human Feedback (RLHF) roles?
What is the difference between Rlhf vs Rn?
| Aspect | Rlhf | Rn |
|---|---|---|
| Required Credentials | Licensed healthcare professional, often with specialized training in mental health or behavioral health | Licensed practical nurse or registered nurse, with nursing licensure and possibly additional certifications |
| Work Environment | Behavioral health facilities, clinics, hospitals, or community health settings | Hospitals, clinics, long-term care facilities, and community health settings |
| Employer & Industry Usage | Behavioral health and mental health services | General healthcare and nursing services |
| Common Search & Comparison | Rlhf vs Rn | Rlhf vs Rn |
While Rlhf (Registered Licensed Mental Health Facilitator) focuses on mental health support and behavioral health interventions, Rn (Registered Nurse) provides broader nursing care across various medical settings. Both roles require licensure, but Rlhf specializes in mental health, whereas Rn covers general patient care.
What are the most commonly searched types of Rlhf jobs in Washington?
The most popular types of Rlhf jobs in Washington are:
What are popular job titles related to Rlhf jobs in Washington?
For Rlhf jobs in Washington, the most frequently searched job titles are:
What job categories do people searching Rlhf jobs in Washington look for?
The top searched job categories for Rlhf jobs in Washington are:
What cities in Washington are hiring for Rlhf jobs?
Cities in Washington with the most Rlhf job openings:

Full-time
Re-posted 23 days ago
Job description
-
Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch
-
Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration
-
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues
-
Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy
-
Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams
-
Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics
-
Implement methods from recent ML papers quickly and turn them into production-grade systems