1

Reinforcement Learning With Human Feedback Jobs in Colorado

Sr AI/ML Engineer

Lone Tree, CO · On-site

$106K - $146K/yr

... human resources, legal, IT, information security, facilities, marketing, and communications ... Hands-on experience with reinforcement learning and real-time systems applicable to MPC.

Sr AI/ML Engineer

Lone Tree, CO · On-site

$106K - $146K/yr

... human resources, legal, IT, information security, facilities, marketing, and communications ... Hands-on experience with reinforcement learning and real-time systems applicable to MPC.

Sr AI/ML Engineer

Longmont, CO

$103K - $141K/yr

... human resources, legal, IT, information security, facilities, marketing, and communications ... Hands-on experience with reinforcement learning and real-time systems applicable to MPC.

Mentor other data scientists through code review, experiment design, and technical feedback and ... reinforcement learning, and deep learning, with a strong foundation in statistical modeling, A/B ...

Mentor other data scientists through code review, experiment design, and technical feedback and ... reinforcement learning, and deep learning, with a strong foundation in statistical modeling, A/B ...

AI/ML Engineer II

Lone Tree, CO

$99K - $136K/yr

Experience with reinforcement learning or generative AI models (e.g., GANs, Transformers). * Working knowledge of Agile or DevOps practices in software/ML project environments. * Hands-on experience ...

AI/ML Engineer II

Lone Tree, CO · On-site

$99K - $136K/yr

Experience with reinforcement learning or generative AI models (e.g., GANs, Transformers). * Working knowledge of Agile or DevOps practices in software/ML project environments. * Hands-on experience ...

Experience in Reinforcement Learning (RL), Computer Vision, or Natural Language Processing (NLP) is ... Experience with PyTorch, TensorFlow, or other deep learning frameworks is required. An advanced ...

Sr AI/ML Engineer

Englewood, CO · On-site

$143.49 - $197.29/hr

Hands‑on experience with reinforcement learning and real‑time systems applicable to MPC. Qualifications - Prefer * Master's degree or Ph.D. in Artificial Intelligence, Machine Learning, or a ...

Showing results 21-40

Reinforcement Learning With Human Feedback information

What is reinforcement learning with human feedback?

Reinforcement Learning with Human Feedback (RLHF) is a machine learning technique where AI agents are trained not only through automated reward signals but also by incorporating feedback from humans. This approach helps align the agent’s behavior with human preferences, values, or safety requirements by allowing humans to guide or correct the learning process. RLHF is commonly used in developing advanced AI systems, such as language models, to ensure their outputs are helpful, safe, and aligned with user expectations. The process often involves human evaluators ranking or scoring the AI's responses, which are then used to fine-tune the model’s behavior.

What collaborations are typical for a reinforcement learning with human feedback specialist within a machine learning team?

As an RLHF specialist, you often work closely with data scientists, machine learning engineers, and domain experts to design effective feedback mechanisms and reward models. Collaboration with annotation teams or subject matter experts is common, as high-quality human feedback is crucial for training robust RLHF models. You may also partner with product managers and UX researchers to ensure that the models align with user needs and ethical considerations. Regular cross-functional meetings and code reviews help maintain alignment and foster innovation across teams.

What are the key skills and qualifications needed to thrive as a reinforcement learning with human feedback engineer?

To excel as a Reinforcement Learning with Human Feedback (RLHF) Engineer, you need a strong background in machine learning, reinforcement learning theory, statistics, and typically an advanced degree in computer science or a related field. Familiarity with deep learning frameworks (such as TensorFlow or PyTorch), RL libraries (like Ray RLlib), and experience with data collection and annotation systems are essential. Excellent problem-solving abilities, communication skills, and teamwork help you collaborate with researchers, data annotators, and other engineers. These skills enable you to design and implement RLHF systems that are robust, scalable, and aligned with human values.

What is the difference between Reinforcement Learning With Human Feedback vs Reinforcement Learning Engineer?

AspectReinforcement Learning With Human FeedbackReinforcement Learning Engineer
CredentialsTypically requires knowledge of machine learning, AI, and data analysisRequires similar credentials in machine learning, programming, and AI
Work EnvironmentResearch labs, AI development teams, tech companiesDevelopment teams, research labs, tech firms
Industry UsageUsed in AI training, human-in-the-loop systems, and model refinementDesigning, implementing, and optimizing reinforcement learning algorithms

Reinforcement Learning With Human Feedback focuses on improving AI models through human input, while Reinforcement Learning Engineers develop and deploy these algorithms. Both roles require strong machine learning skills and often work in similar environments, but their core responsibilities differ in application and focus.

What are popular job titles related to Reinforcement Learning With Human Feedback jobs in Colorado?

For Reinforcement Learning With Human Feedback jobs in Colorado, the most frequently searched job titles are:

What job categories do people searching Reinforcement Learning With Human Feedback jobs in Colorado look for?

The top searched job categories for Reinforcement Learning With Human Feedback jobs in Colorado are:

What cities in Colorado are hiring for Reinforcement Learning With Human Feedback jobs?

Cities in Colorado with the most Reinforcement Learning With Human Feedback job openings:

Infographic showing various Reinforcement Learning With Human Feedback job openings in Colorado as of August 2026, with employment types broken down into 1% As Needed, 76% Full Time, 21% Part Time, 1% Temporary, and 1% Contract. Highlights an 87% Physical, 2% Hybrid, and 11% Remote job distribution.

Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)

Block

Fort Collins, CO • Remote

Full-time

Posted yesterday

New


Block rating

7.9

Company rating: 7.9 out of 10

Based on 16 frontline employees who took The Breakroom Quiz

9th of 21 rated payment service providers


Job description

Team: Apollo - Block Applied R&D Location: Remote (US / Canada) Duration: Fall/Winter 2026 co-op - 8 months, flexible start September 2026 Level: Graduate student (MS or PhD, returning to your program after the co-op)

About Apollo

Apollo leads Block's efforts to build the Customer World Model (CWM): a continuously evolving representation of each customer's goals, context, history, constraints, and likely future needs.

The CWM powers proactive intelligence across Block's ecosystem. Instead of customers navigating products in search of features, intelligence observes their world, understands what matters, anticipates what comes next, and initiates actions on their behalf.

We believe the next generation of AI products will not be defined by chat interfaces or isolated agents. They will be defined by rich world models that enable systems to reason over a customer's evolving state, make better decisions, and learn continuously from outcomes. Apollo designs, prototypes, and guides the development of this intelligence layer.

About the role

We're hiring a small cohort of graduate research interns to help build the foundations of proactive intelligence.

This is not a traditional internship. You'll own a research problem end-to-end: framing the question, developing methods, running experiments, publishing findings, and, when successful, shipping your work into production systems used by millions of customers and sellers.

You'll work at the intersection of representation learning, foundation models, reinforcement learning, causal reasoning, agentic systems, and product intelligence. The goal is not simply to build smarter models, but to build systems that develop a deeper understanding of customers and use that understanding to make better decisions over time.

Past interns have shipped production systems within months and published their work in the same year.

What you'll work on

Depending on your interests and Apollo's roadmap, you'll focus on one or more of the following areas:

Customer World Models

Building rich representations of customers from event streams, financial activity, operational signals, and behavioral data.

Examples include:

  • Representation learning over long-horizon customer histories
  • Event-based foundation models
  • Multi-modal customer representations spanning structured, sequential, and graph data
  • Memory architectures for long-term customer understanding

Proactive Intelligence

Developing systems that can anticipate customer needs and initiate helpful actions before being asked.

Examples include:

  • Opportunity detection and next-best-action systems
  • Long-horizon planning and decision-making
  • Preference and goal inference
  • Learning when intervention creates value versus friction

Agentic Decision Systems

Building agents that reason over customer world models and take actions in real environments.

Examples include:

  • Tool use and planning
  • Multi-step reasoning over customer state
  • Autonomous workflow execution
  • Recovery and adaptation under uncertainty

Learning from Feedback Loops

Developing methods that allow intelligence to improve continuously from real-world outcomes.

Examples include:

  • Reinforcement learning from customer and product feedback
  • Reward modeling and preference learning
  • Counterfactual evaluation
  • Credit assignment over long decision horizons

Evaluation and Measurement

Building evaluation frameworks that predict real-world performance, trust, and customer value.

Examples include:

  • Simulated customer environments
  • Longitudinal evaluation
  • Decision quality metrics
  • Safety and reliability benchmarks

What we're looking for

We're looking for researchers interested in building systems that understand people, learn from experience, and improve over time.

Required

  • Currently enrolled in an MS or PhD program in Computer Science, Machine Learning, Statistics, Mathematics, Operations Research, or a related field, and returning to that program after the co-op.
  • Strong foundations in modern machine learning, including deep learning, optimization, representation learning, and foundation models.
  • Experience conducting independent research and translating ideas into working systems.
  • Fluency in Python and experience with PyTorch, JAX, or similar frameworks.
  • Evidence of research excellence through publications, open-source contributions, technical leadership, or equivalent work.

Nice to have

  • Experience with large language models and agentic systems.
  • Experience with reinforcement learning, reward modeling, or sequential decision-making.
  • Experience with representation learning for structured, temporal, or graph data.
  • Familiarity with large-scale training and production ML systems.
  • Interest in building AI systems that directly affect customer outcomes.

What you'll get

  • Direct mentorship from researchers working on the future of proactive intelligence at Block.
  • Access to large-scale datasets, modern infrastructure, frontier models, and substantial compute resources.
  • Opportunities to publish and contribute to open-source projects.
  • A chance to shape foundational technology that could power the next generation of Block products.
  • Exposure to both scientific research and product deployment, with a clear path from idea to impact.

Application Guidelines

Candidates may submit up to 9 active applications within a 60-day period. Reapplications to the same role are accepted 90 days after a previous application has been reviewed.

Use of AI in Our Hiring Process

We may use automated AI tools to evaluate job applications for efficiency and consistency. These tools comply with local regulations, including bias audits, and we handle all personal data in accordance with state and local privacy laws.

Contact us here with hiring practice or data usage questions.

Every benefit we offer is designed with one goal: empowering you to do the best work of your career while building the life you want. Remote work, medical insurance, flexible time off, retirement savings plans, and modern family planning are just some of our offering. Check out our other benefits at Block.

Block, Inc. (NYSE: XYZ) builds technology to increase access to the global economy. Each of our brands unlocks different aspects of the economy for more people. Square makes commerce and financial services accessible to sellers. Cash App is the easy way to spend, send, and store money. Afterpay is transforming the way customers manage their spending over time. TIDAL is a music platform that empowers artists to thrive as entrepreneurs. Bitkey is a simple self-custody wallet built for bitcoin. Proto is a suite of bitcoin mining products and services. Together, we're helping build a financial system that is open to everyone.


What Block employees say

Pay

Hours and flexibility

Workplace

Get the full story on Breakroom