AI Research Engineer
Palo Alto, CA ยท On-site
Stay current on LLM agents, RL (offline/online, RLHF/RLAIF), constrained decoding, and program synthesis. What Makes You a Great Fit * PhD in CS/AI/ML (or equivalent research experience) with ...
Palo Alto, CA ยท On-site
Stay current on LLM agents, RL (offline/online, RLHF/RLAIF), constrained decoding, and program synthesis. What Makes You a Great Fit * PhD in CS/AI/ML (or equivalent research experience) with ...
Palo Alto, CA ยท On-site
Stay current on LLM agents, RL (offline/online, RLHF/RLAIF), constrained decoding, and program synthesis. What Makes You a Great Fit * PhD in CS/AI/ML (or equivalent research experience) with ...
Mountain View, CA ยท On-site
$21.50/hr
... online and offline initiatives. โข Experiment with data-led marketing, A/B testing, and content ... Soul AI offers generative AI solutions for enterprises, focusing on RLHF and custom AI services.
Mountain View, CA ยท On-site
$21.50/hr
... online and offline initiatives. โข Experiment with data-led marketing, A/B testing, and content ... Soul AI offers generative AI solutions for enterprises, focusing on RLHF and custom AI services.
... online experimentation. Personalization is a first-class objective on this team. You will build ... RLHF), knowledge distillation, and listwise ranking losses (e.g. LambdaLoss, ListNet, ListMLE ...
... online experimentation. Personalization is a first-class objective on this team. You will build ... RLHF), knowledge distillation, and listwise ranking losses (e.g. LambdaLoss, ListNet, ListMLE ...
... online experimentation. Personalization is a first-class objective on this team. You will build ... RLHF), knowledge distillation, and listwise ranking losses (e.g. LambdaLoss, ListNet, ListMLE ...
... online experimentation. Personalization is a first-class objective on this team. You will build ... RLHF), knowledge distillation, and listwise ranking losses (e.g. LambdaLoss, ListNet, ListMLE ...
Fremont, CA ยท On-site
$150K - $250K/yr
Design and evolve data + evaluation systems inspired by RL from human preferences (RLHF) and ... online RL pipelines). * Depth in deep learning, sequence modeling, and generative models.
Fremont, CA ยท On-site
$150K - $250K/yr
Design and evolve data + evaluation systems inspired by RL from human preferences (RLHF) and ... online RL pipelines). * Depth in deep learning, sequence modeling, and generative models.
... online experimentation. Personalization is a first-class objective on this team. You will build ... GRPO, DPO, RLHF), knowledge distillation, and listwise ranking losses (e.g. LambdaLoss, ListNet ...
... online experimentation. Personalization is a first-class objective on this team. You will build ... GRPO, DPO, RLHF), knowledge distillation, and listwise ranking losses (e.g. LambdaLoss, ListNet ...
... offline/online eval pipelines, agent simulations, data pipelines, backend services, and ... RLHF/RLVR evaluation) - enabling the next generation of AI-powered products, agents, automation ...
... offline/online eval pipelines, agent simulations, data pipelines, backend services, and ... RLHF/RLVR evaluation) - enabling the next generation of AI-powered products, agents, automation ...
... RLHF for longโhorizon agentic tasks. * Dataโdriven decision making: Define AI evaluation as a ... Build online evaluation--A/B testing, production telemetry, user feedback loops, and anomaly ...
... RLHF for longโhorizon agentic tasks. * Dataโdriven decision making: Define AI evaluation as a ... Build online evaluation--A/B testing, production telemetry, user feedback loops, and anomaly ...
... online learning and recommendation systems Experience working with machine learning or LLM model ... techniques (e.g RLHF, Reward model, DPO, PPO, GRPO etc.), Parameter efficient fine-tuning ...
... online learning and recommendation systems Experience working with machine learning or LLM model ... techniques (e.g RLHF, Reward model, DPO, PPO, GRPO etc.), Parameter efficient fine-tuning ...
$122K - $168K/yr
... RL/RLHF/RLAIF) Ensure train/serve consistency - that the inference path used in RL and evaluation faithfully matches production online behavior (e.g., parity across sampling and logit processing ...
$122K - $168K/yr
... RL/RLHF/RLAIF) Ensure train/serve consistency - that the inference path used in RL and evaluation faithfully matches production online behavior (e.g., parity across sampling and logit processing ...
$122K - $168K/yr
... RL/RLHF/RLAIF) Ensure train/serve consistency - that the inference path used in RL and evaluation faithfully matches production online behavior (e.g., parity across sampling and logit processing ...
$122K - $168K/yr
... RL/RLHF/RLAIF) Ensure train/serve consistency - that the inference path used in RL and evaluation faithfully matches production online behavior (e.g., parity across sampling and logit processing ...
$150K - $250K/yr
Design and evolve data + evaluation systems inspired by RL from human preferences (RLHF) and ... online RL pipelines). * Depth in deep learning, sequence modeling, and generative models.
$150K - $250K/yr
Design and evolve data + evaluation systems inspired by RL from human preferences (RLHF) and ... online RL pipelines). * Depth in deep learning, sequence modeling, and generative models.
Cupertino, CA ยท On-site
$143K - $286K/yr
Define experimentation strategies, offline evaluation frameworks, and online/offline metric ... or RLHF. * Strong engineering skills in Python and Scala, with experience building large-scale ...
Cupertino, CA ยท On-site
$143K - $286K/yr
Define experimentation strategies, offline evaluation frameworks, and online/offline metric ... or RLHF. * Strong engineering skills in Python and Scala, with experience building large-scale ...
Fremont, CA ยท On-site
$143K - $286K/yr
Define experimentation strategies, offline evaluation frameworks, and online/offline metric ... or RLHF. * Strong engineering skills in Python and Scala, with experience building large-scale ...
Fremont, CA ยท On-site
$143K - $286K/yr
Define experimentation strategies, offline evaluation frameworks, and online/offline metric ... or RLHF. * Strong engineering skills in Python and Scala, with experience building large-scale ...
Sunnyvale, CA ยท On-site
$122K - $168K/yr
... RL/RLHF/RLAIF) Ensure train/serve consistency - that the inference path used in RL and evaluation faithfully matches production online behavior (e.g., parity across sampling and logit processing ...
Sunnyvale, CA ยท On-site
$122K - $168K/yr
... RL/RLHF/RLAIF) Ensure train/serve consistency - that the inference path used in RL and evaluation faithfully matches production online behavior (e.g., parity across sampling and logit processing ...
... and RLHF for long-horizon agentic tasks. With the aim of improving overall tools performance ... Define and build online evaluation: A/B testing, production telemetry, user feedback loops, and ...
... and RLHF for long-horizon agentic tasks. With the aim of improving overall tools performance ... Define and build online evaluation: A/B testing, production telemetry, user feedback loops, and ...
Redwood City, CA ยท On-site
$125K - $165K/yr
Lead AI agent evaluation systems, including offline and online evaluation pipelines, golden ... Experience with LLM fine-tuning pipelines (SFT, RLHF/RLAIF, preference learning, domain adaptation)
Redwood City, CA ยท On-site
$125K - $165K/yr
Lead AI agent evaluation systems, including offline and online evaluation pipelines, golden ... Experience with LLM fine-tuning pipelines (SFT, RLHF/RLAIF, preference learning, domain adaptation)
Sunnyvale, CA ยท On-site
$122K - $168K/yr
... RL/RLHF/RLAIF) Ensure train/serve consistency - that the inference path used in RL and evaluation faithfully matches production online behavior (e.g., parity across sampling and logit processing ...
Sunnyvale, CA ยท On-site
$122K - $168K/yr
... RL/RLHF/RLAIF) Ensure train/serve consistency - that the inference path used in RL and evaluation faithfully matches production online behavior (e.g., parity across sampling and logit processing ...
San Mateo, CA ยท On-site
$143K - $286K/yr
Define experimentation strategies, offline evaluation frameworks, and online/offline metric ... or RLHF. * Strong engineering skills in Python and Scala, with experience building large-scale ...
San Mateo, CA ยท On-site
$143K - $286K/yr
Define experimentation strategies, offline evaluation frameworks, and online/offline metric ... or RLHF. * Strong engineering skills in Python and Scala, with experience building large-scale ...
... RLHF for long-horizon agentic tasks. With the aim of improving overall tools performance. Data ... Define and build online evaluation: A/B testing, production telemetry, user feedback loops, and ...
... RLHF for long-horizon agentic tasks. With the aim of improving overall tools performance. Data ... Define and build online evaluation: A/B testing, production telemetry, user feedback loops, and ...
| Aspect | Online Rlhf | Online Rlhf |
|---|---|---|
| Credentials | Typically requires certification in online health coaching or related fields | Typically requires certification in online health coaching or related fields |
| Work Environment | Remote, online platform-based | Remote, online platform-based |
| Industry Usage | Common in health and wellness sectors | Common in health and wellness sectors |
| Job Focus | Providing health guidance and support online | Providing health guidance and support online |
Online Rlhf and Online Rlhf are the same role, often used interchangeably. Both involve providing health and wellness support remotely, requiring similar certifications and working within the online health industry. The key difference is often in terminology rather than job function.
The most popular types of Rlhf jobs in California are:
For Online Rlhf jobs in California, the most frequently searched job titles are:
The top searched job categories for Online Rlhf jobs in California are:
Cities in California with the most Online Rlhf job openings:

Palo Alto, CA โข On-site
Other
Re-posted 8 days ago
Normal Computing builds silicon that turns thermal noise from an obstacle into a computational resource. Conventional chips spend most of their energy forcing determinism onto physics; ours compute with it. Stochastic, in-memory, asynchronous: the result is 10-100ร more AI inference per dollar, per watt.
We co-design the full stack: AI-native EDA systems in production with the world's largest semiconductor companies, and the advanced ASICs they make possible. Backed by $85M+ from the world's leading deep-tech investors and built by scientists, engineers, and operators from the labs that built modern computing.
Normal works as one team across New York, Silicon Valley, London, Copenhagen, and Seoul. We hire people who want the hardest version of their craft, across every discipline, at every seniority.
Your Role in Our MissionWeโre hiring an AI Research Engineer to push the frontier of agentic LLMs and reinforcement learning for our agentic code generation tool. Youโll design and run experiments, build agents, curate datasets from complex technical documents (e.g., chip specifications), and create rigorous evaluations. Youโll write productionโquality research code and work closely with engineering to ship improvements to customers. Leadership not requiredโimpact through research and building is.
ResponsibilitiesDesign and implement multiโagent and RL approaches for agentic code generation and toolโuse.
Build research prototypes that integrate with our agentic code generation tool; collaborate to productionize wins.
Create evaluation suites: task specs, pass/fail checkers, coverage, cost/latency dashboards.
Acquire and curate datasets from PDFs/logs/tables; generate synthetic data where appropriate; maintain data cards and licensing.
Analyze experiments with disciplined ablations; document results and decisions.
Stay current on LLM agents, RL (offline/online, RLHF/RLAIF), constrained decoding, and program synthesis.
PhD in CS/AI/ML (or equivalent research experience) with publications ideally in multiโagent RL, agentic AI, or RL for language/code.
Strong Python and ML framework experience (PyTorch preferred; JAX/HF a plus).
Demonstrated ability to turn research into working systems; reproducibility mindset (tests, seeds, configs, logging).
Experience designing eval harnesses and success metrics for sequential/agentic tasks.
Comfortable with data acquisition/curation from documents/logs; good instincts about data quality and licenses.
Clear communicator who partners well with engineers.
Research on program synthesis/codegen, constrained decoding, or executionโbased rewards.
Experience with offline RL from tool traces or human corrections.
Openโsource contributions (e.g., CleanRL, RLlib, AutoGen, LangGraph, CrewAI, Transformers).
Familiarity with semiconductor/chip domains or other complex technical specs.
Track record of shipping research to production and measuring impact.
Equal Employment Opportunity Statement
Normal Computing is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, or any other legally protected status.
Accessibility Accommodations
Normal Computing is committed to providing reasonable accommodations to individuals with disabilities. If you need assistance or an accommodation due to a disability, please let us know at accommodations@normalcomputing.com.
Privacy Notice
By submitting your application, you agree that Normal Computing may collect, use, and store your personal information for employment-related purposes in accordance with our Privacy Policy.