Our partner is looking for a Research Engineer (Reinforcement Learning) based in Netherlands. Join a small, senior engineering team building the next generation of voice- and text-driven AI agents.
Our partner is looking for a Research Engineer (Reinforcement Learning) based in Netherlands. Join a small, senior engineering team building the next generation of voice- and text-driven AI agents.
Staff, Data Scientist (Pricing/Reinforcement Learning)
Noel, MO · On-site
$110K - $220K/yr
This is a full-time onsite role based in Bentonville, AR or Sunnyvale, CA; remote and hybrid ... Optimization & Reinforcement Learning: Multi-armed bandits, Deep RL (PPO, DQN) for sequential ...
Staff, Data Scientist (Pricing/Reinforcement Learning)
Noel, MO · On-site
$110K - $220K/yr
This is a full-time onsite role based in Bentonville, AR or Sunnyvale, CA; remote and hybrid ... Optimization & Reinforcement Learning: Multi-armed bandits, Deep RL (PPO, DQN) for sequential ...
Staff, Data Scientist (Pricing/Reinforcement Learning)
Cassville, MO · On-site
$110K - $220K/yr
This is a full-time onsite role based in Bentonville, AR or Sunnyvale, CA; remote and hybrid ... Optimization & Reinforcement Learning: Multi-armed bandits, Deep RL (PPO, DQN) for sequential ...
Staff, Data Scientist (Pricing/Reinforcement Learning)
Cassville, MO · On-site
$110K - $220K/yr
This is a full-time onsite role based in Bentonville, AR or Sunnyvale, CA; remote and hybrid ... Optimization & Reinforcement Learning: Multi-armed bandits, Deep RL (PPO, DQN) for sequential ...
Staff, Data Scientist (Pricing/Reinforcement Learning)
Anderson, MO · On-site
$110K - $220K/yr
This is a full-time onsite role based in Bentonville, AR or Sunnyvale, CA; remote and hybrid ... Optimization & Reinforcement Learning: Multi-armed bandits, Deep RL (PPO, DQN) for sequential ...
Staff, Data Scientist (Pricing/Reinforcement Learning)
Anderson, MO · On-site
$110K - $220K/yr
This is a full-time onsite role based in Bentonville, AR or Sunnyvale, CA; remote and hybrid ... Optimization & Reinforcement Learning: Multi-armed bandits, Deep RL (PPO, DQN) for sequential ...
AI Motion Intern - VinMotion US
California, MO · On-site
$90 - $120/hr
Technical Skills Proficiency in Python, C++, TensorFlow, PyTorch. Experience with Reinforcement Learning (RL), Model Predictive Control (MPC), Whole‑Body Control (WBC), and Imitation Learning.
New
AI Motion Intern - VinMotion US
California, MO · On-site
$90 - $120/hr
Technical Skills Proficiency in Python, C++, TensorFlow, PyTorch. Experience with Reinforcement Learning (RL), Model Predictive Control (MPC), Whole‑Body Control (WBC), and Imitation Learning.
New
(USA) Principal, Machine Learning Engineer
California, MO · On-site
$169 - $338/hr
Drive innovation in applied ML, including deep learning, NLP, reinforcement learning, and multimodal modeling approaches. * Mentor engineers, influence architectural decisions, and promote ...
New
(USA) Principal, Machine Learning Engineer
California, MO · On-site
$169 - $338/hr
Drive innovation in applied ML, including deep learning, NLP, reinforcement learning, and multimodal modeling approaches. * Mentor engineers, influence architectural decisions, and promote ...
New
Postdoctoral Research Associate - Oncology
Saint Louis, MO · On-site
$55 - $70/hr
David Spencer's Lab in the Division of Oncology, Department of Medicine. Dr. David Spencer is ... deep learning methods. There will be opportunities for development of new cell line and animal ...
Postdoctoral Research Associate - Oncology
Saint Louis, MO · On-site
$55 - $70/hr
David Spencer's Lab in the Division of Oncology, Department of Medicine. Dr. David Spencer is ... deep learning methods. There will be opportunities for development of new cell line and animal ...
Postdoctoral Research Associate - Engineering
Saint Louis, MO · On-site
$55 - $65/hr
## Postdoctoral Research Associate - EngineeringApplyremote type: On Campus/Onsitelocations ... Strong programming skills and experience with machine learning applications in geoscience are ...
Postdoctoral Research Associate - Engineering
Saint Louis, MO · On-site
$55 - $65/hr
## Postdoctoral Research Associate - EngineeringApplyremote type: On Campus/Onsitelocations ... Strong programming skills and experience with machine learning applications in geoscience are ...
The postdoc will also have the opportunity to work on other topics related to air quality modeling ... Strong programming skills and experience with machine learning applications in geoscience are ...
The postdoc will also have the opportunity to work on other topics related to air quality modeling ... Strong programming skills and experience with machine learning applications in geoscience are ...
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Posted today
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Posted today
Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)
Columbia, MO · On-site
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)
Columbia, MO · On-site
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)
Kansas City, MO · On-site
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)
Kansas City, MO · On-site
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)
Columbia, MO · Remote
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)
Columbia, MO · Remote
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)
Independence, MO · On-site
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)
Independence, MO · On-site
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)
Independence, MO · Remote
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)
Independence, MO · Remote
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)
Jefferson City, MO · Remote
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)
Jefferson City, MO · Remote
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)
Saint Louis, MO · On-site
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)
Saint Louis, MO · On-site
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Posted today
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Posted today
The postdoc will also have the opportunity to work on other topics related to air quality modeling ... Strong programming skills and experience with machine learning applications in geoscience are ...
The postdoc will also have the opportunity to work on other topics related to air quality modeling ... Strong programming skills and experience with machine learning applications in geoscience are ...
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Posted today
You'll work at the intersection of representation learning, foundation models, reinforcement ... Strong foundations in modern machine learning, including deep learning, optimization ...
Posted today
Postdoctoral In Reinforcement Learning information
What is a postdoctoral researcher in reinforcement learning?
What are the key skills and qualifications needed to thrive as a postdoctoral researcher in reinforcement learning?
What are some common challenges faced by postdoctoral researchers in reinforcement learning, and how can they be addressed?
What is the difference between Postdoctoral In Reinforcement Learning vs Postdoctoral In Machine Learning?
| Aspect | Postdoctoral In Reinforcement Learning | Postdoctoral In Machine Learning |
|---|---|---|
| Required Credentials | PhD in Computer Science, AI, or related field; strong programming skills; research experience in reinforcement learning | PhD in Computer Science, AI, or related field; strong programming skills; research experience in machine learning |
| Work Environment | Academic labs, research institutions, industry R&D teams focused on reinforcement learning applications | Academic labs, research institutions, industry R&D teams working on various machine learning techniques |
| Industry Usage | Primarily in AI research, robotics, gaming, and autonomous systems | Broader applications including data analysis, predictive modeling, and AI research |
Postdoctoral In Reinforcement Learning specializes in research related to decision-making algorithms and autonomous systems, whereas Postdoctoral In Machine Learning covers a wider range of AI techniques. Both roles require similar credentials but differ in focus and application areas.
What are popular job titles related to Postdoctoral In Reinforcement Learning jobs in Missouri?
For Postdoctoral In Reinforcement Learning jobs in Missouri, the most frequently searched job titles are:
What job categories do people searching Postdoctoral In Reinforcement Learning jobs in Missouri look for?
The top searched job categories for Postdoctoral In Reinforcement Learning jobs in Missouri are:
What cities in Missouri are hiring for Postdoctoral In Reinforcement Learning jobs?
Cities in Missouri with the most Postdoctoral In Reinforcement Learning job openings:

Full-time
Medical, Dental, Vision, PTO
This job post has expired today. Applications are no longer accepted.
Job description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Research Engineer (Reinforcement Learning) based in Netherlands.
Join a small, senior engineering team building the next generation of voice- and text-driven AI agents.
You'll focus on post-training models to make agents more capable, reliable, and effective over long-running interactions.
Your work will span environments, verifiers, synthetic data, training experiments, evaluations, and production deployment.
You'll tackle challenging problems such as persistent context, reliable tool use, and multi-turn agent behavior.
The role combines hands-on research and engineering, with a strong emphasis on measurable improvements in model performance.
You'll work closely with experienced engineers in a remote, collaborative environment where technical craft and creativity are highly valued.
Your contributions will directly shape AI systems operating at significant production scale.
- Build training environments, verifiers, and supporting infrastructure for post-training models.
- Own the synthetic data pipeline from data generation through quality assurance and validation.
- Run end-to-end training experiments, analyze results, and clearly identify the factors driving model improvements.
- Design and maintain evaluations that models must pass before production releases.
- Select and adapt suitable open-weight foundation models for specific agent and product requirements.
- Develop trained behaviors that perform consistently across both voice and text-based agents.
- Deploy trained models to production and continuously improve them based on real-world usage and feedback.
- Develop robust approaches to long-horizon interactions, accumulated context, and reliable tool use during live conversations.
- Strong Python engineering skills and the ability to build reliable, production-quality systems.
- Demonstrated experience taking a machine learning model from raw data through experimentation and into production.
- A strong data-centric mindset, with attention to coverage, diversity, quality, and data leakage.
- The ability to anticipate reward exploitation and design robust rewards, verifiers, and evaluation mechanisms.
- Practical experience working with GPUs and a realistic understanding of their capabilities and limitations.
- Strong judgment around when model training is the right solution-and when a simpler approach is preferable.
- Ability to collaborate effectively within a remote, distributed, and highly autonomous team.
- Experience with post-training techniques such as fine-tuning, reward design, or reinforcement learning, including approaches such as GRPO, is highly desirable.
- Familiarity with RL and fine-tuning frameworks such as TRL, verl, OpenRLHF, or custom training loops is a plus.
- Experience with technologies such as vLLM or SGLang for fast rollouts and FSDP for multi-GPU training is advantageous.
- Experience training tool-using or multi-turn agents, as well as building execution sandboxes, verifiers, evaluation harnesses, or developer tooling, is valuable.
- Familiarity with open-weight model families such as Qwen or Llama and techniques such as LoRA is a plus.
- Opportunity to make a significant impact on a fast-growing developer platform and help shape its future.
- Collaboration with a small, highly experienced team that values technical excellence, creativity, and ownership.
- Competitive salary and equity package.
- Health, dental, and vision benefits.
- Flexible vacation policy.
- Remote-friendly working environment with flexibility and autonomy.