San Francisco, CA · On-site · Full-time Compensation: $180,000–$220,000 + competitive equity (profit-share brings total cash ~$500K)
About the Company
Our client builds the training data and evaluation infrastructure that frontier AI labs use to improve their models, designing high-signal datasets and running rigorous evaluations that go beyond static benchmarks. It's a small, early team where individual contributors have direct impact on how the next generation of models learns. The founding team comes from top quant firms, big tech, and leading AI research labs.
Founded 2025 · 11–50 people · Industry: AI / ML
The Role
As a Software Engineer (Environments), you design the datasets and evaluation rubrics that directly influence how frontier models learn — going from hypothesis to live experiment quickly, with output feeding directly into model training.
What you'll be doing
- Design targeted data slices that surface model failure modes across high-stakes domains (finance, code generation, enterprise workflows)
- Build and iterate on evaluation rubrics and reward signals powering RLHF and RLVR training pipelines
- Develop quantitative frameworks to measure dataset quality, diversity, and downstream impact on model alignment and capability
- Own end-to-end real-world and synthetic data pipelines, from scoping with research teams to production-ready evaluation specs
- Run annotator-modeling experiments to improve model capabilities across task types
Requirements
- 1–4 years of software-engineering experience with strong technical depth
Nice to Haves
- Experience at RL-environment companies
- A background in AI safety or benchmarking organizations (e.g., METR, Artificial Analysis)
- A genuine obsession with how data structure, selection, and quality drive model behavior
- An ability to design lightweight experiments and move fast
- Former founders or early engineers at early-stage startups
- A demonstrated ability to work hard, learn fast, and care deeply about details
Why Join
- Outsized total cash: base plus profit share brings expected total cash to around $500K, plus competitive equity
- A direct line to frontier models: output feeds directly into model training runs at scale, working hands-on with research teams at top AI labs
- High ownership on an early team: scope, build, and ship end to end
Details
- Location — San Francisco, CA
- Work policy — On-site
- Compensation — $180,000–$220,000 + competitive equity (profit-share brings total cash ~$500K)
- Visa sponsorship — Available (O-1, OPT)
- Employment type — Full-time