1

Postdoctoral In Reinforcement Learning Jobs in Alameda, CA

Senior Staff AI Engineer

San Francisco, CA · On-site

$65 - $84/hr

About the Role We are seeking an experienced AI Engineer with deep expertise in Reinforcement Learning (RL) to join our team as a Senior Staff Architect. In this role, you will be responsible for ...

Showing results 21-40

Postdoctoral In Reinforcement Learning information

See Alameda, CA salary details

$28.3K

$66.9K

$94.6K

How much do postdoctoral in reinforcement learning jobs pay per year?

As of Aug 21, 2026, the average yearly pay for postdoctoral in reinforcement learning in Alameda, CA is $66,893.00, according to ZipRecruiter salary data. Most workers in this role earn between $55,500.00 and $75,400.00 per year, depending on experience, location, and employer.

What is a postdoctoral researcher in reinforcement learning?

A Postdoctoral Researcher in Reinforcement Learning is an individual who has completed a PhD and conducts advanced research in the field of reinforcement learning, a branch of artificial intelligence focused on how agents take actions in environments to maximize rewards. These researchers often work in academic, industrial, or governmental research settings, collaborating on projects that advance the theoretical foundations or practical applications of reinforcement learning. Their responsibilities may include designing experiments, developing algorithms, publishing papers, and mentoring graduate students.

What are the key skills and qualifications needed to thrive as a postdoctoral researcher in reinforcement learning?

To thrive as a Postdoctoral Researcher in Reinforcement Learning, you need a PhD in computer science or a related field, with deep expertise in machine learning, statistics, and algorithm development. Proficiency in programming languages such as Python, experience with deep learning frameworks (e.g., TensorFlow or PyTorch), and familiarity with reinforcement learning libraries are typically required. Strong analytical thinking, problem-solving ability, collaboration, and scientific communication skills help you excel in research teams and publish impactful work. These competencies are vital to advancing state-of-the-art research, developing novel algorithms, and contributing to the academic and industrial progress in AI.

What are some common challenges faced by postdoctoral researchers in reinforcement learning, and how can they be addressed?

Postdoctoral researchers in reinforcement learning often face challenges such as balancing independent research projects with collaborative work, staying up-to-date with rapidly evolving literature, and managing the pressure to publish in top conferences. Effective time management, regular engagement with the research community through seminars and workshops, and seeking mentorship from senior colleagues can help address these challenges. Additionally, collaborating with interdisciplinary teams can offer fresh perspectives and support, making it easier to navigate complex research problems.

What is the difference between Postdoctoral In Reinforcement Learning vs Postdoctoral In Machine Learning?

AspectPostdoctoral In Reinforcement LearningPostdoctoral In Machine Learning
Required CredentialsPhD in Computer Science, AI, or related field; strong programming skills; research experience in reinforcement learningPhD in Computer Science, AI, or related field; strong programming skills; research experience in machine learning
Work EnvironmentAcademic labs, research institutions, industry R&D teams focused on reinforcement learning applicationsAcademic labs, research institutions, industry R&D teams working on various machine learning techniques
Industry UsagePrimarily in AI research, robotics, gaming, and autonomous systemsBroader applications including data analysis, predictive modeling, and AI research

Postdoctoral In Reinforcement Learning specializes in research related to decision-making algorithms and autonomous systems, whereas Postdoctoral In Machine Learning covers a wider range of AI techniques. Both roles require similar credentials but differ in focus and application areas.

What cities near Alameda, CA are hiring for Postdoctoral In Reinforcement Learning jobs?

Cities near Alameda, CA with the most Postdoctoral In Reinforcement Learning job openings:

Infographic showing various Postdoctoral In Reinforcement Learning job openings in Alameda, CA as of August 2026, with employment types broken down into 1% As Needed, 73% Full Time, 25% Part Time, and 1% Contract. Highlights an 87% Physical, 2% Hybrid, and 11% Remote job distribution, with an average salary of $66,893 per year, or $32.2 per hour.

Research Scientist - Post-training / RL

Epsilon Labs, Inc.

San Francisco, CA • On-site

$180 - $240/hr

Other

Posted 3 days ago

New


Job description

About Us

We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting‑edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product‑market fit with a substantial customer pipeline already in place.

Role Overview

We're seeking a Research Scientist with deep expertise in post‑training and reinforcement learning to join our ML Research team. You'll be at the forefront of developing and deploying state‑of‑the‑art multimodal models for clinical use in radiology settings. This role owns every stage after pretraining: supervised fine‑tuning, reward modeling, reinforcement learning against verifiable and learned reward signals, reasoning and tool‑use training, and inference‑time strategy. You'll work with one of the largest and most diverse medical imaging datasets in the industry, advancing the state‑of‑the‑art in grounded report generation, reward design, and inference‑time reasoning while maintaining the clinical rigor required for healthcare deployment.

Key Responsibilities
  • Design reinforcement learning with verifiable rewards for report generation, including clinical label and entity‑relation matching, grounding IoU, measurement accuracy, and reporting schema compliance.

  • Extend reinforcement learning to unverifiable and noisy objectives such as report quality and clinical usefulness, using learned reward models and radiologist feedback pipelines (RLHF) built on expert preferences and report edits.

  • Run GRPO‑family algorithms with complex multi‑reward objectives, tuning reward composition and diagnosing reward hacking, entropy collapse, and diversity loss.

  • Train explicit reward models, including multimodal reward models conditioned on the image, with both outcome and process supervision.

  • Train chain‑of‑thought reasoning over image regions, including evidence localization and verification loops that keep reasoning grounded in the image rather than in language priors.

  • Train multimodal tool use — windowing, zoom and crop, detector and segmentation calls, prior study retrieval — with credit assignment across multi‑turn trajectories.

  • Develop inference‑time methods including best‑of‑N sampling against reward models and grounding‑aware decoding, and distill the resulting gains back into the policy.

  • Tune output stylization to institutional reporting conventions, keeping style rewards separated from clinical content rewards.

  • Stay current with cutting‑edge research in reinforcement learning, reward modeling, and multimodal post‑training.

  • Drive research and technical excellence through conference publications and technical blog posts, establishing best practices for post‑training medical VLMs at scale.

Qualifications
  • 6+ years of academia/industry experience in reinforcement learning, post‑training, or multimodal machine learning

  • Deep expertise in post‑training large language or vision‑language models (e.g., Qwen‑VL, InternVL, LLaVA, or similar architectures)

  • Strong foundation in modern post‑training and reinforcement learning techniques including:

    • Group‑relative policy optimization and its successors (GRPO, DAPO, GSPO, CISPO) with multi‑reward objectives

    • Reinforcement learning with verifiable rewards, and with noisy, sparse, or learned reward signals

    • Reward model training: pairwise and generative reward models, outcome and process supervision

    • Preference optimization methods (DPO, IPO, ORPO, KTO) and RLHF

    • Inference‑time compute scaling, including best‑of‑N sampling and verifier‑guided decoding

  • Practical experience diagnosing and mitigating reward hacking and reward over‑optimization

  • Track record of implementing complex models from research papers and adapting them to new domains

  • Proficiency in PyTorch or JAX, with experience training large models on multi‑GPU/distributed systems

  • Experience with reinforcement learning infrastructure at scale, including rollout generation (vLLM, SGLang) and frameworks such as verl, TRL, or OpenRLHF

  • Experience with autoregressive language modeling and instruction tuning

  • Strong software engineering skills and ability to write production‑quality code

Preferred Qualifications
  • Publications at top‑tier conferences (NeurIPS, ICML, ICLR, CVPR, ACL, EMNLP, MICCAI)

  • Hands‑on experience with medical imaging applications, particularly radiology report generation

  • Experience with agentic or multi‑turn reinforcement learning, including credit assignment over tool‑use trajectories

  • Experience with grounded generation tasks (visual grounding, referring expression comprehension)

  • Knowledge of evaluation methodologies for long‑form generation, including factuality assessment and hallucination detection

  • Experience mitigating catastrophic forgetting of supervised capabilities during reinforcement learning

  • Familiarity with clinical NLP and medical knowledge representation

  • Experience with model interpretability, explainability, and uncertainty quantification in safety‑critical applications

#J-18808-Ljbffr