1

Online Rlhf Jobs in California (NOW HIRING)

LLM Training Engineer

San Francisco, CA · On-site

$155K - $220K/yr

Design offline + online environments that support RL-style training at scale * Instrument ... Post-training pipelines (SFT, RLHF/RLAIF, preference optimization, eval loops) * Building RL ...

AI Marketer/DevRel

Mountain View, CA · On-site

$21.50/hr

... online and offline initiatives. • Experiment with data-led marketing, A/B testing, and content ... Soul AI offers generative AI solutions for enterprises, focusing on RLHF and custom AI services.

Senior AI Engineer

Los Angeles, CA · On-site

$112K - $154K/yr

End-to-End ML Lifecycle:  Own requirements → data prep → feature engineering → classical ML or LLM fine-tuning (LoRA, PEFT, RLHF) → offline/online evaluation → MLflow registry, with ...

Senior AI Engineer

Los Angeles, CA · On-site

$112K - $154K/yr

Own requirements data prep feature engineering classical ML or LLM fine-tuning (LoRA, PEFT, RLHF) offline/online evaluation MLflow registry, with automated drift and quality alerts. * Data & Storage ...

Senior Inference Engineer, AGI

Sunnyvale, CA · On-site

$122K - $168K/yr

... RL/RLHF/RLAIF) Ensure train/serve consistency - that the inference path used in RL and evaluation faithfully matches production online behavior (e.g., parity across sampling and logit processing ...

Define experimentation strategies, offline evaluation frameworks, and online/offline metric ... or RLHF. * Strong engineering skills in Python and Scala, with experience building large-scale ...

Define experimentation strategies, offline evaluation frameworks, and online/offline metric ... or RLHF. * Strong engineering skills in Python and Scala, with experience building large-scale ...

next page

Showing results 1-20

Online Rlhf information

What is an online RLHF?

Online RLHF (Reinforcement Learning from Human Feedback) jobs typically involve helping to train AI models by providing human feedback on their outputs. Workers in these roles might review model responses, rate the quality of generated text, or suggest improvements to help the AI learn to produce better results. These jobs are often remote and can be done part-time or as contract work. They play a crucial role in improving the safety, usefulness, and accuracy of AI systems by aligning them more closely with human preferences.

What are some common challenges faced by online RLHF specialists when collaborating with cross-functional teams?

Online RLHF specialists often work closely with machine learning engineers, data annotators, and product managers. A common challenge is ensuring that feedback from human annotators is accurately integrated into model training, which requires clear communication and well-defined annotation guidelines. Additionally, balancing the pace of model updates with the need for high-quality human feedback can be demanding. Effective collaboration and regular syncs are essential to maintain alignment and achieve project goals.

What are the key skills and qualifications needed to thrive as an online RLHF specialist, and why are they important?

To thrive as an Online RLHF Specialist, you need a strong background in machine learning, reinforcement learning, and data analysis, typically supported by a degree in computer science or a related field. Familiarity with technical tools like Python, PyTorch or TensorFlow, and experience with human feedback systems or annotation platforms are highly valuable. Strong problem-solving, attention to detail, and the ability to communicate complex concepts clearly are crucial soft skills. These qualifications ensure the effective training and evaluation of AI models, leading to more accurate and reliable machine learning systems.

What is the difference between Online Rlhf vs Online Rlhf?

AspectOnline RlhfOnline Rlhf
CredentialsTypically requires certification in online health coaching or related fieldsTypically requires certification in online health coaching or related fields
Work EnvironmentRemote, online platform-basedRemote, online platform-based
Industry UsageCommon in health and wellness sectorsCommon in health and wellness sectors
Job FocusProviding health guidance and support onlineProviding health guidance and support online

Online Rlhf and Online Rlhf are the same role, often used interchangeably. Both involve providing health and wellness support remotely, requiring similar certifications and working within the online health industry. The key difference is often in terminology rather than job function.

What are the most commonly searched types of Rlhf jobs in California?

The most popular types of Rlhf jobs in California are:

What job categories do people searching Online Rlhf jobs in California look for?

The top searched job categories for Online Rlhf jobs in California are:

What cities in California are hiring for Online Rlhf jobs?

Cities in California with the most Online Rlhf job openings:

Infographic showing various Online Rlhf job openings in California as of August 2026, with employment types broken down into 68% Full Time, 29% Part Time, 1% Temporary, and 2% Contract. Highlights an 80% Physical, 1% Hybrid, and 19% Remote job distribution.

$150 - $210/hr

Other

Posted 16 days ago


Job description

About Normal Computing

Normal Computing builds silicon that turns thermal noise from an obstacle into a computational resource. Conventional chips spend most of their energy forcing determinism onto physics; ours compute with it. Stochastic, in-memory, asynchronous: the result is 10-100× more AI inference per dollar, per watt.

We co-design the full stack: AI-native EDA systems in production with the world's largest semiconductor companies, and the advanced ASICs they make possible. Backed by $85M+ from the world's leading deep-tech investors and built by scientists, engineers, and operators from the labs that built modern computing.

Normal works as one team across New York, Silicon Valley, London, Copenhagen, and Seoul. We hire people who want the hardest version of their craft, across every discipline, at every seniority.

Your Role in Our Mission

We’re hiring an AI Research Engineer to push the frontier of agentic LLMs and reinforcement learning for our agentic code generation tool. You’ll design and run experiments, build agents, curate datasets from complex technical documents (e.g., chip specifications), and create rigorous evaluations. You’ll write production‑quality research code and work closely with engineering to ship improvements to customers. Leadership not required—impact through research and building is.

Responsibilities
  • Design and implement multi‑agent and RL approaches for agentic code generation and tool‑use.

  • Build research prototypes that integrate with our agentic code generation tool; collaborate to productionize wins.

  • Create evaluation suites: task specs, pass/fail checkers, coverage, cost/latency dashboards.

  • Acquire and curate datasets from PDFs/logs/tables; generate synthetic data where appropriate; maintain data cards and licensing.

  • Analyze experiments with disciplined ablations; document results and decisions.

  • Stay current on LLM agents, RL (offline/online, RLHF/RLAIF), constrained decoding, and program synthesis.

What Makes You a Great Fit
  • PhD in CS/AI/ML (or equivalent research experience) with publications ideally in multi‑agent RL, agentic AI, or RL for language/code.

  • Strong Python and ML framework experience (PyTorch preferred; JAX/HF a plus).

  • Demonstrated ability to turn research into working systems; reproducibility mindset (tests, seeds, configs, logging).

  • Experience designing eval harnesses and success metrics for sequential/agentic tasks.

  • Comfortable with data acquisition/curation from documents/logs; good instincts about data quality and licenses.

  • Clear communicator who partners well with engineers.

Bonus Points For
  • Research on program synthesis/codegen, constrained decoding, or execution‑based rewards.

  • Experience with offline RL from tool traces or human corrections.

  • Open‑source contributions (e.g., CleanRL, RLlib, AutoGen, LangGraph, CrewAI, Transformers).

  • Familiarity with semiconductor/chip domains or other complex technical specs.

  • Track record of shipping research to production and measuring impact.

Equal Employment Opportunity Statement

Normal Computing is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, or any other legally protected status.

Accessibility Accommodations

Normal Computing is committed to providing reasonable accommodations to individuals with disabilities. If you need assistance or an accommodation due to a disability, please let us know at accommodations@normalcomputing.com.

Privacy Notice

By submitting your application, you agree that Normal Computing may collect, use, and store your personal information for employment-related purposes in accordance with our Privacy Policy.

#J-18808-Ljbffr