Bespoke Labs

60 Bespoke Labs Jobs Hiring Near You

* Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Build and ...

* Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Shape environments end to end: stateful, resumable ...

* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...

* Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Build and ...

* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...

* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...

* Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts * Shape environments end to end: stateful, resumable ...

* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...

Showing results 41-60

Bespoke Labs Jobs Information

Infographic showing various job openings at Bespoke Labs in the United States as of August 2026, with employment types broken down into 100% Full Time. Highlights an 100% Physical job distribution.

Long-Horizon Coding Task Expert

Bespoke Labs

Murfreesboro, TN

Full-time

Posted 5 days ago


Job description

  • Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts

  • Shape environments end to end: stateful, resumable systems with snapshotting, checkpointing, and branching rollouts, at multi-node scale where needed

  • Generate and refine ideas for tasks and agentic trajectories across multi-turn, tool-using, and computer-use agent loops with persistent state across turns

  • Design reward structure for sparse-reward settings: milestone and process rewards, subgoal and task decomposition, and credit assignment across long trajectories

  • Validate that the work is correct, hard, and covers the right edge cases, including rubric design for partial credit, contamination avoidance, and defenses against reward hacking

  • Design curricula that ramp task difficulty rather than shipping fixed-difficulty tasks

  • Build automated pipelines that curate high-quality long-horizon RL environments and tasks at scale

  • Write reliable, well-tested Python infrastructure rather than one-off research scripts

  • Work directly with our research team with high ownership, building novel systems rather than maintaining legacy ones