Machine Learning Engineer
Kennesaw, GA ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Kennesaw, GA ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Kennesaw, GA ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Atlanta, GA ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Atlanta, GA ยท On-site
SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment ...
Fine-tuning, instruction tuning, RLHF, and domain adaptation * Model Evaluation: Performance metrics, benchmark development, and A/B testing frameworks * Consulting mindset with a strong bias for ...
Fine-tuning, instruction tuning, RLHF, and domain adaptation * Model Evaluation: Performance metrics, benchmark development, and A/B testing frameworks * Consulting mindset with a strong bias for ...
Apply techniques such as prompt engineering, RAG (Retrieval-Augmented Generation), fine-tuning, and RLHF to enhance model performance. * Develop and deploy autonomous AI agents using frameworks like ...
Apply techniques such as prompt engineering, RAG (Retrieval-Augmented Generation), fine-tuning, and RLHF to enhance model performance. * Develop and deploy autonomous AI agents using frameworks like ...
Alpharetta, GA ยท On-site
Apply techniques such as prompt engineering, RAG (Retrieval-Augmented Generation), fine-tuning, and RLHF to enhance model performance. * Develop and deploy autonomous AI agents using frameworks like ...
Alpharetta, GA ยท On-site
Apply techniques such as prompt engineering, RAG (Retrieval-Augmented Generation), fine-tuning, and RLHF to enhance model performance. * Develop and deploy autonomous AI agents using frameworks like ...
Atlanta, GA ยท Remote
$100 - $200/hr
Familiarity with prompt engineering, AI output evaluation, fact-checking, or RLHF is a plus. * Master's, MBA, PhD, or other advanced degree is preferred.
Quick apply
Atlanta, GA ยท Remote
$100 - $200/hr
Familiarity with prompt engineering, AI output evaluation, fact-checking, or RLHF is a plus. * Master's, MBA, PhD, or other advanced degree is preferred.
Atlanta, GA ยท Remote
$100 - $200/hr
Familiarity with prompt engineering, AI output evaluation, fact-checking, or RLHF is a plus. * Master's, MBA, PhD, or other advanced degree is preferred.
Quick apply
Atlanta, GA ยท Remote
$100 - $200/hr
Familiarity with prompt engineering, AI output evaluation, fact-checking, or RLHF is a plus. * Master's, MBA, PhD, or other advanced degree is preferred.
Atlanta, GA ยท Remote
$100 - $200/hr
Familiarity with prompt engineering, AI output evaluation, fact-checking, or RLHF is a plus. * Master's, MBA, PhD, or other advanced degree is preferred.
New
Quick apply
Atlanta, GA ยท Remote
$100 - $200/hr
Familiarity with prompt engineering, AI output evaluation, fact-checking, or RLHF is a plus. * Master's, MBA, PhD, or other advanced degree is preferred.
New
Atlanta, GA ยท On-site
Advance Workday's proprietary capabilities in pre-training, post-training (RLHF, DPO), and domain-specific alignment for HR and Finance workflows. * Publish & Open Source: Lead Workday's contribution ...
Atlanta, GA ยท On-site
Advance Workday's proprietary capabilities in pre-training, post-training (RLHF, DPO), and domain-specific alignment for HR and Finance workflows. * Publish & Open Source: Lead Workday's contribution ...
Atlanta, GA ยท On-site
$150K - $215K/yr
Exposure to dataset curation and post-training techniques (SFT, DPO, RLHF) on open-weight models using tools like Axolotl, Unsloth, Hugging Face transformers, or TRL. * Experience with trace mining ...
Atlanta, GA ยท On-site
$150K - $215K/yr
Exposure to dataset curation and post-training techniques (SFT, DPO, RLHF) on open-weight models using tools like Axolotl, Unsloth, Hugging Face transformers, or TRL. * Experience with trace mining ...
Specialized expertise in other topics like fine-tuning, RLHF, RAG and Knowledge graph etc. * Experience in designing and implementing Model Context Protocol (MCP) servers to enable seamless ...
Specialized expertise in other topics like fine-tuning, RLHF, RAG and Knowledge graph etc. * Experience in designing and implementing Model Context Protocol (MCP) servers to enable seamless ...
Specialized expertise in other topics like fine-tuning, RLHF, RAG and Knowledge graph etc. * Experience in designing and implementing Model Context Protocol (MCP) servers to enable seamless ...
Specialized expertise in other topics like fine-tuning, RLHF, RAG and Knowledge graph etc. * Experience in designing and implementing Model Context Protocol (MCP) servers to enable seamless ...
Atlanta, GA ยท On-site
SFT, RLHF, PPO, DPO, GRPO, RLAIF, and preference optimization workflows. โข Experience with open-weight foundation models such as Llama, Qwen, Mistral, DeepSeek, or equivalent architectures. โข ...
Atlanta, GA ยท On-site
SFT, RLHF, PPO, DPO, GRPO, RLAIF, and preference optimization workflows. โข Experience with open-weight foundation models such as Llama, Qwen, Mistral, DeepSeek, or equivalent architectures. โข ...
Atlanta, GA ยท On-site
Build and optimize training using techniques such as SFT, RLHF, PPO, DPO, GRPO, RLAIF, and Constitutional AI, and understand how each affects reasoning quality, safety, latency, cost, and reliability.
Atlanta, GA ยท On-site
Build and optimize training using techniques such as SFT, RLHF, PPO, DPO, GRPO, RLAIF, and Constitutional AI, and understand how each affects reasoning quality, safety, latency, cost, and reliability.
Atlanta, GA ยท On-site +1
Working knowledge of model fine-tuning approaches (LoRA, RLHF, instruction tuning, or equivalent) * Experience designing and implementing custom LLM and agent evaluation frameworks, including ...
Atlanta, GA ยท On-site +1
Working knowledge of model fine-tuning approaches (LoRA, RLHF, instruction tuning, or equivalent) * Experience designing and implementing custom LLM and agent evaluation frameworks, including ...
Atlanta, GA ยท On-site +1
Working knowledge of model fine-tuning approaches (LoRA, RLHF, instruction tuning, or equivalent) * Experience designing and implementing custom LLM and agent evaluation frameworks, including ...
Atlanta, GA ยท On-site +1
Working knowledge of model fine-tuning approaches (LoRA, RLHF, instruction tuning, or equivalent) * Experience designing and implementing custom LLM and agent evaluation frameworks, including ...
Atlanta, GA ยท On-site +1
Working knowledge of model fine-tuning approaches (LoRA, RLHF, instruction tuning, or equivalent) * Experience designing and implementing custom LLM and agent evaluation frameworks, including ...
Atlanta, GA ยท On-site +1
Working knowledge of model fine-tuning approaches (LoRA, RLHF, instruction tuning, or equivalent) * Experience designing and implementing custom LLM and agent evaluation frameworks, including ...
... RLHF, RLAIF, fine-tuning, or distillation to smaller models) into production Published or open-source work in agent infrastructure or evaluation tooling Cost management (FinOps) experience for LLM ...
... RLHF, RLAIF, fine-tuning, or distillation to smaller models) into production Published or open-source work in agent infrastructure or evaluation tooling Cost management (FinOps) experience for LLM ...
An RLHF (Reinforcement Learning with Human Feedback) job involves training AI models using human feedback to improve their responses. Professionals in this role analyze model outputs, provide evaluations, and refine AI behavior through reinforcement learning techniques. These roles are common in AI research, content moderation, and chatbot development.
| Aspect | Rlhf | Rn |
|---|---|---|
| Required Credentials | Licensed healthcare professional, often with specialized training in mental health or behavioral health | Licensed practical nurse or registered nurse, with nursing licensure and possibly additional certifications |
| Work Environment | Behavioral health facilities, clinics, hospitals, or community health settings | Hospitals, clinics, long-term care facilities, and community health settings |
| Employer & Industry Usage | Behavioral health and mental health services | General healthcare and nursing services |
| Common Search & Comparison | Rlhf vs Rn | Rlhf vs Rn |
While Rlhf (Registered Licensed Mental Health Facilitator) focuses on mental health support and behavioral health interventions, Rn (Registered Nurse) provides broader nursing care across various medical settings. Both roles require licensure, but Rlhf specializes in mental health, whereas Rn covers general patient care.
The top searched job categories for Rlhf jobs in Decatur, GA are:
Cities near Decatur, GA with the most Rlhf job openings:
Kennesaw, GA โข On-site
Full-time
Re-posted 28 days ago
Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch
Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues
Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy
Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams
Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics
Implement methods from recent ML papers quickly and turn them into production-grade systems