2

Remote Reinforcement Learning Jobs (NOW HIRING)

Machine Learning Engineer

Washington, DC · On-site +1

$130K - $200K/yr

Familiarity or experience with model distillation, synthetic data generation, reinforcement ... Fully remote, U.S.-based * Health Benefits : Comprehensive health, dental, and vision coverage

... reinforcement learning, LLM orchestration, RAG systems. * 3+ years of experience with modern NLP tools and machine learning libraries (scikit-learn, PyTorch, TensorFlow, spaCy) * Experience with unit ...

We use reinforcement learning algorithms to provide this intelligence, converting raw sensor data ... We are a 100% remote company. * Competitive compensation & meaningful equity. * Outsized ...

Data Scientist

San Francisco, CA · On-site +1

$160K - $200K/yr

... Reinforcement Learning, Statistics, and Optimization. The role will report directly to the CTO. ... This is a remote position, but we do have an office in San Fransisco. You will be the first data ...

Expression of Interest - Engineering

Seattle, WA · Remote

$19 - $24.75/hr

We use reinforcement learning algorithms to provide this intelligence, converting raw sensor data ... We are a 100% remote company. * Competitive compensation & meaningful equity. * Outsized ...

Data Scientist

San Francisco, CA · Remote

$160K - $200K/yr

... Reinforcement Learning, Statistics, and Optimization. The role will report directly to the CTO. ... This is a remote position, but we do have an office in San Fransisco. You will be the first data ...

... reinforcement learning, LLM orchestration, RAG systems. * 5+ years of experience with modern NLP ... City, NY (remote), and Seattle, WA (remote). Candidates must permanently reside in the US ...

next page

Showing results 1-20

Remote Reinforcement Learning information

See salary details

$11K

$83.9K

$140K

How much do remote reinforcement learning jobs pay per year?

As of Jul 23, 2026, the average yearly pay for remote reinforcement learning in the United States is $83,885.00, according to ZipRecruiter salary data. Most workers in this role earn between $72,000.00 and $139,000.00 per year, depending on experience, location, and employer.

What is a Remote Reinforcement Learning job?

A Remote Reinforcement Learning job involves developing and applying reinforcement learning algorithms while working from a location outside of a traditional office environment. Professionals in this field focus on creating systems where agents learn optimal behaviors through trial and error, often using feedback from their environment. These jobs typically require expertise in machine learning, programming, and mathematics, and are commonly found in industries like robotics, gaming, and autonomous systems. Working remotely allows researchers and engineers to collaborate with global teams using digital tools and platforms.

What are the key skills and qualifications needed to thrive as a Remote Reinforcement Learning Engineer, and why are they important?

To thrive as a Remote Reinforcement Learning Engineer, you need a strong background in machine learning, statistics, and programming (especially Python), often supported by an advanced degree in computer science or a related field. Familiarity with frameworks such as TensorFlow, PyTorch, and RL-specific libraries like OpenAI Gym, along with experience using cloud computing platforms, is typically required. Excellent problem-solving skills, self-motivation, and effective remote communication help individuals excel in distributed teams. These skills ensure the successful design, implementation, and deployment of reinforcement learning solutions while collaborating efficiently in a remote work environment.

What is the difference between Remote Reinforcement Learning vs Remote Machine Learning Engineer?

AspectRemote Reinforcement Learning
Required CredentialsMaster's or PhD in Computer Science, AI, or related fields; knowledge of RL algorithms
Work EnvironmentResearch-focused, experimental, often involves simulation and algorithm development
Employer & Industry UsageTech companies, research labs, AI startups focusing on autonomous systems
Common Search & Comparison IntentUnderstanding specialized AI roles, research focus, and technical skills

Remote Reinforcement Learning specialists focus on developing algorithms that enable machines to learn through trial and error in simulated or real environments. In contrast, Remote Machine Learning Engineers typically work on deploying and optimizing various machine learning models across applications. While both roles require strong programming skills and knowledge of AI, reinforcement learning emphasizes decision-making processes, whereas machine learning engineering covers a broader range of models and deployment strategies.

What are common challenges faced when working remotely in a Reinforcement Learning role and how can they be addressed?

Working remotely in a Reinforcement Learning role often involves overcoming communication barriers with cross-functional teams, managing large-scale experiments without on-site resources, and staying updated with rapidly evolving research. To address these challenges, it's important to establish regular check-ins with colleagues, utilize cloud-based platforms for experiment management, and participate in virtual seminars or journal clubs. Developing strong self-motivation and time management skills is also crucial to maintain productivity in a remote environment.
More about Remote Reinforcement Learning jobs
What cities are hiring for Remote Reinforcement Learning jobs? Cities with the most Remote Reinforcement Learning job openings:
What are the most commonly searched types of Reinforcement Learning jobs? The most popular types of Reinforcement Learning jobs are:
What states have the most Remote Reinforcement Learning jobs? States with the most job openings for Remote Reinforcement Learning jobs include:
What job categories do people searching Remote Reinforcement Learning jobs look for? The top searched job categories for Remote Reinforcement Learning jobs are:
Infographic showing various Remote Reinforcement Learning job openings in the United States as of July 2026, with employment types broken down into 70% Full Time, 18% Part Time, and 12% Contract. Highlights an 100% Remote job distribution, with an average salary of $83,885 per year, or $40.3 per hour.
Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)

Applied Research Intern, Proactive Intelligence & Customer World Models (PhD / Graduate Co-op)

Block

Bodega Bay, CA • Remote

Other

Posted 14 days ago


Block rating

7.9

Company rating: 7.9 out of 10

Based on 16 frontline employees who took The Breakroom Quiz

10th of 21 rated payment service providers


Job description

Team: Apollo - Block Applied R&D
Location: Remote (US / Canada)
Duration: Fall/Winter 2026 co-op - 8 months, flexible start September 2026
Level: Graduate student (MS or PhD, returning to your program after the co-op)

About Apollo

Apollo leads Block's efforts to build the Customer World Model (CWM): a continuously evolving representation of each customer's goals, context, history, constraints, and likely future needs.

The CWM powers proactive intelligence across Block's ecosystem. Instead of customers navigating products in search of features, intelligence observes their world, understands what matters, anticipates what comes next, and initiates actions on their behalf.

We believe the next generation of AI products will not be defined by chat interfaces or isolated agents. They will be defined by rich world models that enable systems to reason over a customer's evolving state, make better decisions, and learn continuously from outcomes. Apollo designs, prototypes, and guides the development of this intelligence layer.

About the role

We're hiring a small cohort of graduate research interns to help build the foundations of proactive intelligence.

This is not a traditional internship. You'll own a research problem end-to-end: framing the question, developing methods, running experiments, publishing findings, and, when successful, shipping your work into production systems used by millions of customers and sellers.

You'll work at the intersection of representation learning, foundation models, reinforcement learning, causal reasoning, agentic systems, and product intelligence. The goal is not simply to build smarter models, but to build systems that develop a deeper understanding of customers and use that understanding to make better decisions over time.

Past interns have shipped production systems within months and published their work in the same year.

What you'll work on

Depending on your interests and Apollo's roadmap, you'll focus on one or more of the following areas:

Customer World Models

Building rich representations of customers from event streams, financial activity, operational signals, and behavioral data.

Examples include:

  • Representation learning over long-horizon customer histories
  • Event-based foundation models
  • Multi-modal customer representations spanning structured, sequential, and graph data
  • Memory architectures for long-term customer understanding

Proactive Intelligence

Developing systems that can anticipate customer needs and initiate helpful actions before being asked.

Examples include:

  • Opportunity detection and next-best-action systems
  • Long-horizon planning and decision-making
  • Preference and goal inference
  • Learning when intervention creates value versus friction

Agentic Decision Systems

Building agents that reason over customer world models and take actions in real environments.

Examples include:

  • Tool use and planning
  • Multi-step reasoning over customer state
  • Autonomous workflow execution
  • Recovery and adaptation under uncertainty

Learning from Feedback Loops

Developing methods that allow intelligence to improve continuously from real-world outcomes.

Examples include:

  • Reinforcement learning from customer and product feedback
  • Reward modeling and preference learning
  • Counterfactual evaluation
  • Credit assignment over long decision horizons

Evaluation and Measurement

Building evaluation frameworks that predict real-world performance, trust, and customer value.

Examples include:

  • Simulated customer environments
  • Longitudinal evaluation
  • Decision quality metrics
  • Safety and reliability benchmarks
What we're looking for

We're looking for researchers interested in building systems that understand people, learn from experience, and improve over time.

Required

  • Currently enrolled in an MS or PhD program in Computer Science, Machine Learning, Statistics, Mathematics, Operations Research, or a related field, and returning to that program after the co-op.
  • Strong foundations in modern machine learning, including deep learning, optimization, representation learning, and foundation models.
  • Experience conducting independent research and translating ideas into working systems.
  • Fluency in Python and experience with PyTorch, JAX, or similar frameworks.
  • Evidence of research excellence through publications, open-source contributions, technical leadership, or equivalent work.

Nice to have

  • Experience with large language models and agentic systems.
  • Experience with reinforcement learning, reward modeling, or sequential decision-making.
  • Experience with representation learning for structured, temporal, or graph data.
  • Familiarity with large-scale training and production ML systems.
  • Interest in building AI systems that directly affect customer outcomes.
What you'll get
  • Direct mentorship from researchers working on the future of proactive intelligence at Block.
  • Access to large-scale datasets, modern infrastructure, frontier models, and substantial compute resources.
  • Opportunities to publish and contribute to open-source projects.
  • A chance to shape foundational technology that could power the next generation of Block products.
  • Exposure to both scientific research and product deployment, with a clear path from idea to impact.

What Block employees say

Pay

Hours and flexibility

Workplace

Get the full story on Breakroom