Machine Learning Engineer
Scottsdale, AZ ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Scottsdale, AZ ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Scottsdale, AZ ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Phoenix, AZ ยท On-site
$128K - $176K/yr
You will use advanced techniques like prompt engineering, RLHF, and instruction tuning to ensure our models produce high-quality, context-aware responses. Responsibilities: * Lead the fine-tuning ...
Phoenix, AZ ยท On-site
$128K - $176K/yr
You will use advanced techniques like prompt engineering, RLHF, and instruction tuning to ensure our models produce high-quality, context-aware responses. Responsibilities: * Lead the fine-tuning ...
Mesa, AZ ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Mesa, AZ ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Glendale, AZ ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Glendale, AZ ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Chandler, AZ ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Chandler, AZ ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Flagstaff, AZ ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Flagstaff, AZ ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Phoenix, AZ ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Phoenix, AZ ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Phoenix, AZ ยท On-site
$128K - $176K/yr
You will use advanced techniques like prompt engineering, RLHF, and instruction tuning to ensure our models produce high-quality, context-aware responses. Responsibilities: * Lead the fine-tuning ...
Phoenix, AZ ยท On-site
$128K - $176K/yr
You will use advanced techniques like prompt engineering, RLHF, and instruction tuning to ensure our models produce high-quality, context-aware responses. Responsibilities: * Lead the fine-tuning ...
Phoenix, AZ ยท On-site
$128K - $176K/yr
You will use advanced techniques like prompt engineering, RLHF, and instruction tuning to ensure our models produce high-quality, context-aware responses. Responsibilities: * Lead the fine-tuning ...
Phoenix, AZ ยท On-site
$128K - $176K/yr
You will use advanced techniques like prompt engineering, RLHF, and instruction tuning to ensure our models produce high-quality, context-aware responses. Responsibilities: * Lead the fine-tuning ...
Tucson, AZ ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Tucson, AZ ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Phoenix, AZ ยท Remote
$100 - $200/hr
Familiarity with prompt engineering, AI output evaluation, fact-checking, or RLHF is a plus. * Master's, MBA, PhD, or other advanced degree is preferred.
Quick apply
Phoenix, AZ ยท Remote
$100 - $200/hr
Familiarity with prompt engineering, AI output evaluation, fact-checking, or RLHF is a plus. * Master's, MBA, PhD, or other advanced degree is preferred.
Phoenix, AZ ยท Remote
$100 - $200/hr
Familiarity with prompt engineering, AI output evaluation, fact-checking, or RLHF is a plus. * Master's, MBA, PhD, or other advanced degree is preferred.
Quick apply
Phoenix, AZ ยท Remote
$100 - $200/hr
Familiarity with prompt engineering, AI output evaluation, fact-checking, or RLHF is a plus. * Master's, MBA, PhD, or other advanced degree is preferred.
Experience with foundational models, RLHF, and multi-agent systems. * Interpersonal Skills: Proven track record managing multi-disciplinary teams in a fast-paced environment. Excellent communication ...
Experience with foundational models, RLHF, and multi-agent systems. * Interpersonal Skills: Proven track record managing multi-disciplinary teams in a fast-paced environment. Excellent communication ...
An RLHF (Reinforcement Learning with Human Feedback) job involves training AI models using human feedback to improve their responses. Professionals in this role analyze model outputs, provide evaluations, and refine AI behavior through reinforcement learning techniques. These roles are common in AI research, content moderation, and chatbot development.
| Aspect | Rlhf | Rn |
|---|---|---|
| Required Credentials | Licensed healthcare professional, often with specialized training in mental health or behavioral health | Licensed practical nurse or registered nurse, with nursing licensure and possibly additional certifications |
| Work Environment | Behavioral health facilities, clinics, hospitals, or community health settings | Hospitals, clinics, long-term care facilities, and community health settings |
| Employer & Industry Usage | Behavioral health and mental health services | General healthcare and nursing services |
| Common Search & Comparison | Rlhf vs Rn | Rlhf vs Rn |
While Rlhf (Registered Licensed Mental Health Facilitator) focuses on mental health support and behavioral health interventions, Rn (Registered Nurse) provides broader nursing care across various medical settings. Both roles require licensure, but Rlhf specializes in mental health, whereas Rn covers general patient care.
The most popular types of Rlhf jobs in Arizona are:
For Rlhf jobs in Arizona, the most frequently searched job titles are:
The top searched job categories for Rlhf jobs in Arizona are:
Cities in Arizona with the most Rlhf job openings:

Full-time
Re-posted 17 days ago
Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch
Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues
Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy
Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams
Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics
Implement methods from recent ML papers quickly and turn them into production-grade systems