1

Reinforcement Learning With Human Feedback Jobs in Michigan

Machine Learning Engineer

Dearborn, MI · On-site

$105K - $126K/yr

... fine-tuning (SFT), reinforcement learning from human feedback (RLHF), and instruction tuning ... We have been doing this with pride, dedication and plain, old-fashioned hard work for 24 years

... with real data to enhance operational efficiency. Responsibilities : • Run reinforcement learning experiments in our physically realistic simulators of mineral processing operations, and help turn ...

Machine Learning Engineer

Ann Arbor, MI · On-site

$120K - $180K/yr

Solid grounding in machine learning fundamentals, with working knowledge of modern deep learning; exposure to reinforcement learning is a strong plus. * Proficiency in Python and comfort reading and ...

This role partners closely with leadership and HR Partners to identify development needs, build ... Track program effectiveness using metrics, feedback, and performance outcomes Define Career Pathing ...

... with you! Join our HR Talent Community! At HR Collaborative we are dedicated to Making Work Better ... Enjoy free professional events and learning opportunities. * Hassle-Free Consulting: Focus on what ...

Learning & Development (Open Enrollment, Leadership Program, Talent Assessment, etc.) * Assist with ... HR Projects & Analytics * Assist with HR metrics, reporting, and workforce data analysis via ...

next page

Showing results 1-20

Reinforcement Learning With Human Feedback information

What are the key skills and qualifications needed to thrive as a reinforcement learning with human feedback engineer?

To excel as a Reinforcement Learning with Human Feedback (RLHF) Engineer, you need a strong background in machine learning, reinforcement learning theory, statistics, and typically an advanced degree in computer science or a related field. Familiarity with deep learning frameworks (such as TensorFlow or PyTorch), RL libraries (like Ray RLlib), and experience with data collection and annotation systems are essential. Excellent problem-solving abilities, communication skills, and teamwork help you collaborate with researchers, data annotators, and other engineers. These skills enable you to design and implement RLHF systems that are robust, scalable, and aligned with human values.

What is the difference between Reinforcement Learning With Human Feedback vs Reinforcement Learning Engineer?

AspectReinforcement Learning With Human FeedbackReinforcement Learning Engineer
CredentialsTypically requires knowledge of machine learning, AI, and data analysisRequires similar credentials in machine learning, programming, and AI
Work EnvironmentResearch labs, AI development teams, tech companiesDevelopment teams, research labs, tech firms
Industry UsageUsed in AI training, human-in-the-loop systems, and model refinementDesigning, implementing, and optimizing reinforcement learning algorithms

Reinforcement Learning With Human Feedback focuses on improving AI models through human input, while Reinforcement Learning Engineers develop and deploy these algorithms. Both roles require strong machine learning skills and often work in similar environments, but their core responsibilities differ in application and focus.

What is reinforcement learning with human feedback?

Reinforcement Learning with Human Feedback (RLHF) is a machine learning technique where AI agents are trained not only through automated reward signals but also by incorporating feedback from humans. This approach helps align the agent’s behavior with human preferences, values, or safety requirements by allowing humans to guide or correct the learning process. RLHF is commonly used in developing advanced AI systems, such as language models, to ensure their outputs are helpful, safe, and aligned with user expectations. The process often involves human evaluators ranking or scoring the AI's responses, which are then used to fine-tune the model’s behavior.

What collaborations are typical for a reinforcement learning with human feedback specialist within a machine learning team?

As an RLHF specialist, you often work closely with data scientists, machine learning engineers, and domain experts to design effective feedback mechanisms and reward models. Collaboration with annotation teams or subject matter experts is common, as high-quality human feedback is crucial for training robust RLHF models. You may also partner with product managers and UX researchers to ensure that the models align with user needs and ethical considerations. Regular cross-functional meetings and code reviews help maintain alignment and foster innovation across teams.
What are popular job titles related to Reinforcement Learning With Human Feedback jobs in Michigan? For Reinforcement Learning With Human Feedback jobs in Michigan, the most frequently searched job titles are:
What job categories do people searching Reinforcement Learning With Human Feedback jobs in Michigan look for? The top searched job categories for Reinforcement Learning With Human Feedback jobs in Michigan are:
What cities in Michigan are hiring for Reinforcement Learning With Human Feedback jobs? Cities in Michigan with the most Reinforcement Learning With Human Feedback job openings:
Infographic showing various Reinforcement Learning With Human Feedback job openings in Michigan as of August 2026, with employment types broken down into 4% Internship, 68% Full Time, 19% Part Time, and 9% Contract. Highlights an 100% In-person job distribution.

Senior Machine Learning Engineer - Learned Planning/Reinforcement Learning

Torc Robotics

Ann Arbor, MI • On-site, Remote

$102K - $140K/yr

Full-time

Medical, Dental, Vision, Life, Retirement, PTO

Re-posted 5 days ago


Job description

About the Company
At Torc, we have always believed that autonomous vehicle technology will transform how we travel, move freight, and do business.
A leader in autonomous driving since 2007, Torc has spent over a decade commercializing our solutions with experienced partners. Now a part of the Daimler family, we are focused solely on developing software for automated trucks to transform how the world moves freight.
Join us and catapult your career with the company that helped pioneer autonomous technology, and the first AV software company with the vision to partner directly with a truck manufacturer.
Meet the Team
As a Senior Machine Learning Engineer - Learned Planner / Reinforcement Learning, you will develop and deploy machine learning models that drive decision-making for autonomous trucks. Working closely with teams across perception, prediction, planning, and safety, you will build learned behavior systems that enable safe, efficient, and human-like driving in real-world freight environments.
This role focuses on owning model development and delivery for scoped problem areas, contributing to architecture decisions, and driving improvements in model performance, reliability, and iteration speed within the autonomy stack.
What You'll Do
  • Design, develop, and deploy learned behavior models using approaches such as reinforcement learning, behavior cloning, and imitation learning
  • Own end-to-end model development for scoped problem areas, from data ingestion and training to evaluation and deployment
  • Write production-quality ML code to support scalable training, evaluation, and inference workflows
  • Analyze model performance, identify failure modes, and iterate to improve robustness and generalization across driving scenarios
  • Contribute to training pipelines, data workflows, and infrastructure, including working with large-scale datasets from simulation, fleet logs, and on-vehicle data
  • Collaborate with simulation, validation, and autonomy teams to test and evaluate learned behavior models across diverse environments
  • Support integration of learned planning models into simulation and validation frameworks, enabling faster iteration and improved coverage
  • Contribute to model architecture discussions and technical decision-making within the team
  • Mentor junior engineers on implementation, experimentation, and best practices

What You'll Need to Succeed
  • Bachelor's degree in Computer Science, Robotics, Electrical Engineering, Machine Learning, or related technical field with 6+ years of industry experience, OR Master's degree with 3+ years OR PhD with 1+ years of experience
  • Experience applying reinforcement learning, imitation learning, or sequence modeling to robotics, autonomous systems, or complex control problems
  • Strong programming skills in Python and PyTorch, with experience writing production-quality ML code
  • Experience training, evaluating, and improving models using large-scale datasets and distributed compute environments
  • Solid understanding of ML architectures used in autonomy systems (e.g., transformers, RNNs, graph neural networks, policy networks)
  • Experience debugging model behavior, analyzing performance metrics, and improving model reliability
  • Ability to translate ambiguous problems into structured ML solutions and deliver results independently
  • Experience collaborating cross-functionally to integrate ML models into larger autonomy systems

Bonus Points:
  • Experience in autonomous driving, robotics, or simulation-based training environments
  • Experience with reinforcement learning frameworks or distributed training systems (e.g., Ray)
  • Experience working with simulation environments, scenario generation, or large-scale behavior datasets
  • Familiarity with vehicle dynamics, motion planning, or multi-agent decision-making systems
  • Experience deploying ML models into production or real-world robotics systems
  • Experience with learned planning systems or policy learning in real-world or simulation environments
  • Experience integrating learned behavior models into validation and V&V workflows
  • Background in multi-agent modeling, driver behavior modeling, or long-horizon decision-making systems

Work Location: For this position, we are open to hiring in either the Ann Arbor, MI OR Blacksburg, VA (U.S.) office work locations in a hybrid capacity. We are also open to hiring Remote in the United States
Perks of Being a Full-time Torc'r
Torc cares about our team members and we strive to provide benefits and resources to support their health, work/life balance, and future. Our culture is collaborative, energetic, and team focused. Torc offers:
  • A competitive compensation package that includes a bonus component and stock options
  • 100% paid medical, dental, and vision premiums for full-time employees
  • 401K plan with a 6% employer match
  • Flexibility in schedule and generous paid vacation (available immediately after start date)
  • Company-wide holiday office closures
  • AD+D and Life Insurance

At Torc, we're committed to building a diverse and inclusive workplace. We celebrate the uniqueness of our Torc'rs and do not discriminate based on race, religion, color, national origin, gender (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, gender expression, age, veteran status, or disabilities.
Even if you don't meet 100% of the qualifications listed for this opportunity, we encourage you to apply.
Our compensation reflects the cost of labor across several geographic markets. Pay is based on a number of factors and may vary depending on job-related knowledge, skills, and experience. Torc's total compensation package will also include our corporate bonus and stock option plan. Dependent on the position offered, sign-on payments, relocation, and other forms of compensation may be provided as part of a total compensation package, in addition to a full range of medical, financial, and/or other benefits.
Job ID: 102603
Hiring Range for Job Opening
US Pay Range
$226,400-$271,700 USD