* Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Shape environments end to end: stateful, resumable ...
60 Bespoke Labs Jobs Hiring Near You
* Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Shape environments end to end: stateful, resumable ...
* Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Build and ...
* Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Build and ...
Long-Horizon Coding Task Expert
Newark, DE · On-site
* Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Shape environments end to end: stateful, resumable ...
Long-Horizon Coding Task Expert
Newark, DE · On-site
* Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Shape environments end to end: stateful, resumable ...
Machine Learning Engineer
Merced, CA · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Machine Learning Engineer
Merced, CA · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
AI Agent Trajectory Annotator and Reviewer
Topeka, KS · On-site
$20 - $30/hr
Type: Contract, hourly Location: Remote Hours: 20-30 per week Pay: $20-30/hour, based on experience and language coverage Start: Immediate ABOUT THE ROLE We evaluate how well advanced AI coding ...
AI Agent Trajectory Annotator and Reviewer
Topeka, KS · On-site
$20 - $30/hr
Type: Contract, hourly Location: Remote Hours: 20-30 per week Pay: $20-30/hour, based on experience and language coverage Start: Immediate ABOUT THE ROLE We evaluate how well advanced AI coding ...
Type: Contract, hourly Location: Remote Hours: 20-30 per week Pay: $20-30/hour, based on experience and language coverage Start: Immediate ABOUT THE ROLE We evaluate how well advanced AI coding ...
Type: Contract, hourly Location: Remote Hours: 20-30 per week Pay: $20-30/hour, based on experience and language coverage Start: Immediate ABOUT THE ROLE We evaluate how well advanced AI coding ...
RL Environment Engineer
Annapolis, MD · On-site
* Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Build and ...
RL Environment Engineer
Annapolis, MD · On-site
* Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Build and ...
* Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Shape environments end to end: stateful, resumable ...
* Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Shape environments end to end: stateful, resumable ...
* Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Shape environments end to end: stateful, resumable ...
* Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Shape environments end to end: stateful, resumable ...
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Machine Learning Engineer
Rexburg, ID · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Machine Learning Engineer
Rexburg, ID · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Long-Horizon Coding Task Expert
Casper, WY · On-site
* Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Shape environments end to end: stateful, resumable ...
Long-Horizon Coding Task Expert
Casper, WY · On-site
* Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Shape environments end to end: stateful, resumable ...
Type: Contract, hourly Location: Remote Hours: 20-30 per week Pay: $20-30/hour, based on experience and language coverage Start: Immediate ABOUT THE ROLE We evaluate how well advanced AI coding ...
Type: Contract, hourly Location: Remote Hours: 20-30 per week Pay: $20-30/hour, based on experience and language coverage Start: Immediate ABOUT THE ROLE We evaluate how well advanced AI coding ...
* Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Shape environments end to end: stateful, resumable ...
* Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Shape environments end to end: stateful, resumable ...
AI Agent Trajectory Annotator and Reviewer
Corpus Christi, TX · On-site
$20 - $30/hr
Type: Contract, hourly Location: Remote Hours: 20-30 per week Pay: $20-30/hour, based on experience and language coverage Start: Immediate ABOUT THE ROLE We evaluate how well advanced AI coding ...
AI Agent Trajectory Annotator and Reviewer
Corpus Christi, TX · On-site
$20 - $30/hr
Type: Contract, hourly Location: Remote Hours: 20-30 per week Pay: $20-30/hour, based on experience and language coverage Start: Immediate ABOUT THE ROLE We evaluate how well advanced AI coding ...
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
* Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Shape environments end to end: stateful, resumable ...
* Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Shape environments end to end: stateful, resumable ...
Machine Learning Engineer
Baltimore, MD · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Machine Learning Engineer
Baltimore, MD · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Bespoke Labs Jobs Information

Full-time
Posted 5 days ago
Job description
-
Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts
-
Shape environments end to end: stateful, resumable systems with snapshotting, checkpointing, and branching rollouts, at multi-node scale where needed
-
Generate and refine ideas for tasks and agentic trajectories across multi-turn, tool-using, and computer-use agent loops with persistent state across turns
-
Design reward structure for sparse-reward settings: milestone and process rewards, subgoal and task decomposition, and credit assignment across long trajectories
-
Validate that the work is correct, hard, and covers the right edge cases, including rubric design for partial credit, contamination avoidance, and defenses against reward hacking
-
Design curricula that ramp task difficulty rather than shipping fixed-difficulty tasks
-
Build automated pipelines that curate high-quality long-horizon RL environments and tasks at scale
-
Write reliable, well-tested Python infrastructure rather than one-off research scripts
-
Work directly with our research team with high ownership, building novel systems rather than maintaining legacy ones