1

Helper Reinforcement Learning Jobs in Missouri (NOW HIRING)

Our partner is looking for a Research Engineer (Reinforcement Learning) based in Netherlands. Join ... Opportunity to make a significant impact on a fast-growing developer platform and help shape its ...

$80K - $110K/yr

You will help build inference and fine-tuning technologies for foundation models spanning language ... Deep understanding of the theoretical foundations of machine learning and reinforcement learning.

$11.50 - $15.50/hr

You will help learners understand how data labeling influences machine learning systems, generative ... fine-tuning, reinforcement learning from human feedback (RLHF), and human feedback for model ...

Post-train LLMs and agents -- supervised fine-tuning and reinforcement learning (RLHF/RLAIF, PPO ... Keep track of developments in the field of Artificial Intelligence and help identify, define, and ...

(USA) Director, Data Science

Cassville, MO · On-site

$130K - $260K/yr

Reinforcement learning * Optimization * Experimentation and A/B testing * Promote best practices in ... We believe we are best equipped to help our associates, customers, and the communities we serve ...

(USA) Director, Data Science

Anderson, MO · On-site

$130K - $260K/yr

Reinforcement learning * Optimization * Experimentation and A/B testing * Promote best practices in ... We believe we are best equipped to help our associates, customers, and the communities we serve ...

(USA) Director, Data Science

Noel, MO · On-site

$130K - $260K/yr

Reinforcement learning * Optimization * Experimentation and A/B testing * Promote best practices in ... We believe we are best equipped to help our associates, customers, and the communities we serve ...

You will help equip sales teams with the knowledge, skills, content, and tools needed to improve ... reinforcement, and measurement. * Create high-impact learning experiences and practical enablement ...

BCABA Tutor

Saint Louis, MO · Remote

$18 - $40/hr

About the Job The Varsity Tutors Live Learning Platform has thousands of students looking for ... Ability to explain reinforcement schedules, functional behavior assessment, and behavior ...

BCABA Tutor

Kansas City, MO · Remote

$18 - $40/hr

About the Job The Varsity Tutors Live Learning Platform has thousands of students looking for ... Ability to explain reinforcement schedules, functional behavior assessment, and behavior ...

BCABA Tutor

Columbia, MO · Remote

$18 - $40/hr

About the Job The Varsity Tutors Live Learning Platform has thousands of students looking for ... Ability to explain reinforcement schedules, functional behavior assessment, and behavior ...

Be Seen First

JOB SUMMARY for Certified Special Education Teacher Under the direction of The Behavior Helper ... Utilizes positive reinforcement and behavior modification to instruct students in socially ...

next page

Showing results 1-20

Helper Reinforcement Learning information

What is the difference between Helper Reinforcement Learning vs Data Scientist?

AspectHelper Reinforcement LearningData Scientist
Required CredentialsDegree in Computer Science, AI, or related fields; knowledge of reinforcement learningDegree in Data Science, Statistics, Computer Science; proficiency in programming and analytics
Work EnvironmentResearch labs, AI development teams, tech companiesBusiness analytics, research, consulting firms, tech companies
Industry UsageAI development, machine learning projectsData analysis, predictive modeling, business insights
Common Search/ComparisonHelper Reinforcement Learning vs Data Scientist

Helper Reinforcement Learning focuses on developing algorithms that enable machines to learn through interactions, often requiring knowledge of reinforcement learning techniques. Data Scientists analyze data to extract insights, build models, and support decision-making. While both roles involve programming and data handling, Helper Reinforcement Learning is more specialized in AI algorithm development, whereas Data Scientists work broadly across data analysis and modeling in various industries.

What are the most commonly searched types of Reinforcement Learning jobs in Missouri?

The most popular types of Reinforcement Learning jobs in Missouri are:

What are popular job titles related to Helper Reinforcement Learning jobs in Missouri?

For Helper Reinforcement Learning jobs in Missouri, the most frequently searched job titles are:

What cities in Missouri are hiring for Helper Reinforcement Learning jobs?

Cities in Missouri with the most Helper Reinforcement Learning job openings:

Infographic showing various Helper Reinforcement Learning job openings in Missouri as of August 2026, with employment types broken down into 1% As Needed, 75% Full Time, 22% Part Time, and 2% Contract. Highlights an 86% Physical, 2% Hybrid, and 12% Remote job distribution.

Research Engineer (Reinforcement Learning)

Jobgether

Remote

Full-time

Medical, Dental, Vision, PTO

Posted 4 days ago


Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Research Engineer (Reinforcement Learning) based in Netherlands.

Join a small, senior engineering team building the next generation of voice- and text-driven AI agents.
You'll focus on post-training models to make agents more capable, reliable, and effective over long-running interactions.
Your work will span environments, verifiers, synthetic data, training experiments, evaluations, and production deployment.
You'll tackle challenging problems such as persistent context, reliable tool use, and multi-turn agent behavior.
The role combines hands-on research and engineering, with a strong emphasis on measurable improvements in model performance.
You'll work closely with experienced engineers in a remote, collaborative environment where technical craft and creativity are highly valued.
Your contributions will directly shape AI systems operating at significant production scale.

Accountabilities
  • Build training environments, verifiers, and supporting infrastructure for post-training models.
  • Own the synthetic data pipeline from data generation through quality assurance and validation.
  • Run end-to-end training experiments, analyze results, and clearly identify the factors driving model improvements.
  • Design and maintain evaluations that models must pass before production releases.
  • Select and adapt suitable open-weight foundation models for specific agent and product requirements.
  • Develop trained behaviors that perform consistently across both voice and text-based agents.
  • Deploy trained models to production and continuously improve them based on real-world usage and feedback.
  • Develop robust approaches to long-horizon interactions, accumulated context, and reliable tool use during live conversations.
Requirements:
  • Strong Python engineering skills and the ability to build reliable, production-quality systems.
  • Demonstrated experience taking a machine learning model from raw data through experimentation and into production.
  • A strong data-centric mindset, with attention to coverage, diversity, quality, and data leakage.
  • The ability to anticipate reward exploitation and design robust rewards, verifiers, and evaluation mechanisms.
  • Practical experience working with GPUs and a realistic understanding of their capabilities and limitations.
  • Strong judgment around when model training is the right solution-and when a simpler approach is preferable.
  • Ability to collaborate effectively within a remote, distributed, and highly autonomous team.
  • Experience with post-training techniques such as fine-tuning, reward design, or reinforcement learning, including approaches such as GRPO, is highly desirable.
  • Familiarity with RL and fine-tuning frameworks such as TRL, verl, OpenRLHF, or custom training loops is a plus.
  • Experience with technologies such as vLLM or SGLang for fast rollouts and FSDP for multi-GPU training is advantageous.
  • Experience training tool-using or multi-turn agents, as well as building execution sandboxes, verifiers, evaluation harnesses, or developer tooling, is valuable.
  • Familiarity with open-weight model families such as Qwen or Llama and techniques such as LoRA is a plus.
Benefits:
  • Opportunity to make a significant impact on a fast-growing developer platform and help shape its future.
  • Collaboration with a small, highly experienced team that values technical excellence, creativity, and ownership.
  • Competitive salary and equity package.
  • Health, dental, and vision benefits.
  • Flexible vacation policy.
  • Remote-friendly working environment with flexibility and autonomy.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
 Why Apply Through Jobgether? 
 
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
 
 
#LI-CL1
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
apply for this job