AI Engineer
Manhattan, NY Β· On-site
Research or applied experience with LLM agents, RL (offline/online, RLHF/RLAIF), constrained decoding, or program synthesis. * Openβsource contributions or publications in AI/ML venues. * Skill in ...
Manhattan, NY Β· On-site
Research or applied experience with LLM agents, RL (offline/online, RLHF/RLAIF), constrained decoding, or program synthesis. * Openβsource contributions or publications in AI/ML venues. * Skill in ...
Manhattan, NY Β· On-site
Research or applied experience with LLM agents, RL (offline/online, RLHF/RLAIF), constrained decoding, or program synthesis. * Openβsource contributions or publications in AI/ML venues. * Skill in ...
New York, NY Β· On-site
$200K - $400K/yr
Research or applied experience with LLM agents, RL (offline/online, RLHF/RLAIF), constrained decoding, or program synthesis * Open-source contributions or publications in AI/ML venues * Skill in ...
New York, NY Β· On-site
$200K - $400K/yr
Research or applied experience with LLM agents, RL (offline/online, RLHF/RLAIF), constrained decoding, or program synthesis * Open-source contributions or publications in AI/ML venues * Skill in ...
New York, NY Β· On-site
$200K - $400K/yr
Stay current on LLM agents, RL (offline/online, RLHF/RLAIF), constrained decoding, and program synthesis. What Makes You a Great Fit * PhD in CS/AI/ML (or equivalent research experience) with ...
New York, NY Β· On-site
$200K - $400K/yr
Stay current on LLM agents, RL (offline/online, RLHF/RLAIF), constrained decoding, and program synthesis. What Makes You a Great Fit * PhD in CS/AI/ML (or equivalent research experience) with ...
Palo Alto, CA Β· On-site
Stay current on LLM agents, RL (offline/online, RLHF/RLAIF), constrained decoding, and program synthesis. What Makes You a Great Fit * PhD in CS/AI/ML (or equivalent research experience) with ...
Palo Alto, CA Β· On-site
Stay current on LLM agents, RL (offline/online, RLHF/RLAIF), constrained decoding, and program synthesis. What Makes You a Great Fit * PhD in CS/AI/ML (or equivalent research experience) with ...
Seattle, WA Β· On-site
$300K - $500K/yr
Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement. * Design the systems abstractions that connect research ...
Seattle, WA Β· On-site
$300K - $500K/yr
Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement. * Design the systems abstractions that connect research ...
Experience with agent evaluation, offline/online experiments, and human feedback loops in production. * Direct experience with RLHF, RLAIF, DPO, PPO, GRPO, or related optimization techniques. * Prior ...
Experience with agent evaluation, offline/online experiments, and human feedback loops in production. * Direct experience with RLHF, RLAIF, DPO, PPO, GRPO, or related optimization techniques. * Prior ...
Seattle, WA Β· On-site
Experience with agent evaluation, offline/online experiments, and human feedback loops in production. * Direct experience with RLHF, RLAIF, DPO, PPO, GRPO, or related optimization techniques. * Prior ...
Seattle, WA Β· On-site
Experience with agent evaluation, offline/online experiments, and human feedback loops in production. * Direct experience with RLHF, RLAIF, DPO, PPO, GRPO, or related optimization techniques. * Prior ...
Seattle, WA Β· On-site
$250K - $350K/yr
Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement. * Design the systems abstractions that connect research ...
Seattle, WA Β· On-site
$250K - $350K/yr
Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement. * Design the systems abstractions that connect research ...
Mountain View, CA Β· On-site
$21.50/hr
... online and offline initiatives. β’ Experiment with data-led marketing, A/B testing, and content ... Soul AI offers generative AI solutions for enterprises, focusing on RLHF and custom AI services.
Mountain View, CA Β· On-site
$21.50/hr
... online and offline initiatives. β’ Experiment with data-led marketing, A/B testing, and content ... Soul AI offers generative AI solutions for enterprises, focusing on RLHF and custom AI services.
Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants. * Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF. * Set up ...
Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants. * Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF. * Set up ...
Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants. * Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF. * Set up ...
Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants. * Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF. * Set up ...
Boston, MA Β· On-site
... RLHF), and ensuring our production AI systems maintain world-class accuracy, safety, and ... online LLM observability, output drift, and semantic regression testing). Lead Responsible AI ...
Boston, MA Β· On-site
... RLHF), and ensuring our production AI systems maintain world-class accuracy, safety, and ... online LLM observability, output drift, and semantic regression testing). Lead Responsible AI ...
... from Retail, Online, and Resellers. These solutions are based on cutting edge enterprise ... RLHF, PPO, GRPO) Demonstrated ability to quickly master emerging AI tools and integrate them into ...
... from Retail, Online, and Resellers. These solutions are based on cutting edge enterprise ... RLHF, PPO, GRPO) Demonstrated ability to quickly master emerging AI tools and integrate them into ...
Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement. * Design the systems abstractions that connect research ...
Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement. * Design the systems abstractions that connect research ...
$181K - $290K/yr
... RLHF, RLAIF, or DPO for multiobjective optimization. * Develop reward models and objective ... online and batch adaptation loops with strong guardrails. * Translate conversational logs ...
$181K - $290K/yr
... RLHF, RLAIF, or DPO for multiobjective optimization. * Develop reward models and objective ... online and batch adaptation loops with strong guardrails. * Translate conversational logs ...
... online experimentation. Personalization is a first-class objective on this team. You will build ... RLHF), knowledge distillation, and listwise ranking losses (e.g. LambdaLoss, ListNet, ListMLE ...
... online experimentation. Personalization is a first-class objective on this team. You will build ... RLHF), knowledge distillation, and listwise ranking losses (e.g. LambdaLoss, ListNet, ListMLE ...
... online experimentation. Personalization is a first-class objective on this team. You will build ... RLHF), knowledge distillation, and listwise ranking losses (e.g. LambdaLoss, ListNet, ListMLE ...
... online experimentation. Personalization is a first-class objective on this team. You will build ... RLHF), knowledge distillation, and listwise ranking losses (e.g. LambdaLoss, ListNet, ListMLE ...
Fremont, CA Β· On-site
$150K - $250K/yr
Design and evolve data + evaluation systems inspired by RL from human preferences (RLHF) and ... online RL pipelines). * Depth in deep learning, sequence modeling, and generative models.
Fremont, CA Β· On-site
$150K - $250K/yr
Design and evolve data + evaluation systems inspired by RL from human preferences (RLHF) and ... online RL pipelines). * Depth in deep learning, sequence modeling, and generative models.
Waltham, MA Β· On-site
$175K - $220K/yr
Research reward modeling, and offline-to-online RL for large multimodal policies * Close the sim-2 ... Experience fine-tuning foundation models with RLHF, DPO, GRPO, or related methods * Familiarity ...
Waltham, MA Β· On-site
$175K - $220K/yr
Research reward modeling, and offline-to-online RL for large multimodal policies * Close the sim-2 ... Experience fine-tuning foundation models with RLHF, DPO, GRPO, or related methods * Familiarity ...
... from Retail, Online, and Resellers. These solutions are based on cutting edge enterprise ... RLHF, PPO, GRPO) Demonstrated ability to quickly master emerging AI tools and integrate them into ...
... from Retail, Online, and Resellers. These solutions are based on cutting edge enterprise ... RLHF, PPO, GRPO) Demonstrated ability to quickly master emerging AI tools and integrate them into ...
$17.5K - $23.7K
12% of jobs
$28.2K is the 25th percentile. Wages below this are outliers.
$23.7K - $30K
18% of jobs
$30K - $36.2K
15% of jobs
The median wage is $37.1K / yr.
$36.2K - $42.4K
33% of jobs
$42.4K - $48.6K
12% of jobs
$48.6K - $54.9K
0% of jobs
$54.9K - $61.1K
0% of jobs
$61.1K - $67.3K
0% of jobs
$67.3K - $73.5K
0% of jobs
$73.5K - $79.8K
0% of jobs
$79.8K - $86K
9% of jobs
$17.5K
$40.6K
$86K
| Aspect | Online Rlhf | Online Rlhf |
|---|---|---|
| Credentials | Typically requires certification in online health coaching or related fields | Typically requires certification in online health coaching or related fields |
| Work Environment | Remote, online platform-based | Remote, online platform-based |
| Industry Usage | Common in health and wellness sectors | Common in health and wellness sectors |
| Job Focus | Providing health guidance and support online | Providing health guidance and support online |
Online Rlhf and Online Rlhf are the same role, often used interchangeably. Both involve providing health and wellness support remotely, requiring similar certifications and working within the online health industry. The key difference is often in terminology rather than job function.
Cities with the most Online Rlhf job openings:
The most popular types of Rlhf jobs are:
States with the most job openings for Online Rlhf jobs include:
The top searched job categories for Online Rlhf jobs are:

Manhattan, NY β’ On-site
Other
Re-posted 14 days ago
Normal Computing | Incredible Opportunities
The Normal Team builds foundational software and hardware that help move technology forward - supporting the semiconductor industry, critical AI infrastructure, and the broader systems that power our world. We work as one team across New York, San Francisco, Copenhagen, Seoul, and London.
Your Role in Our Mission:We are looking for an AI Engineer to build production systems that understand large technical documents - like chip design specifications - and turn them into code. You'll ship real improvements to customers weekly while pushing the boundaries of whatβs possible in AI for hardware through reinforcement learning, agentic coding, and exceptional software engineering.
Responsibilities:Lead end-to-end AI development from initial concept through production deployment and iteration.
Design and implement LLM-powered solutions that extract meaning from complex technical specifications.
Handle multi-modal complexity and explore multiβagent and RL approaches for agentic code generation and toolβuse.
Design strategies to manage latency, output variance, and graceful error handling at scale.
Collaborate with product and engineering teams to embed AI capabilities seamlessly into our platform.
Guide junior engineers and establish best practices for AI development
Previous experience delivering production AI systems involving language models, preferably involving document understanding and/or agentic workflows.
Solid software engineering skills with experience in distributed systems and production-grade code.
Proficiency in Python and modern ML frameworks (PyTorch, Hugging Face, transformers.)
Handsβon experience with prompt engineering, fineβtuning, and deploying large language models.
Ability to wrangle, clean, and preprocess large-scale, heterogeneous datasets.
Understanding of AI safety, bias mitigation, and ethical considerations.
Ability to explain complex AI concepts to both technical and nonβtechnical stakeholders.
Experience deploying AI systems in missionβcritical or highβstakes production environments.
Experience with cloud platforms (AWS, GCP, Azure) for largeβscale AI infrastructure.
Research or applied experience with LLM agents, RL (offline/online, RLHF/RLAIF), constrained decoding, or program synthesis.
Openβsource contributions or publications in AI/ML venues.
Skill in balancing cuttingβedge innovation with production reliability and pragmatism.
Normal Computing is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, or any other legally protected status.
Accessibility AccommodationsNormal Computing is committed to providing reasonable accommodations to individuals with disabilities. If you need assistance or an accommodation due to a disability, please let us know at accommodations@normalcomputing.com.
Privacy NoticeBy submitting your application, you agree that Normal Computing may collect, use, and store your personal information for employmentβrelated purposes in accordance with our Privacy Policy.