1

What Can You Do With A Phd In Rhetoric Jobs (NOW HIRING)

Algorithmic Engineer Location: in MD - 15 mins. from DC What you'll get to do: * Research, design ... Prior experience with HE or SMPC is a big plus, but is not required. * Either: * A PhD in Computer ...

What Can We Offer You? * A schedule based on YOUR availability in YOUR city -- we're everywhere ... What Can You Do For Our Clients? * Help them stay in their homes * Some need us to provide personal ...

In this hybrid role you will report to a Data Science Manager. You will: * Develop evaluation ... Collaborate with Product and Engineering partners developing the Waymo Driver and Waymo ...

We're on a quest to empower live communities, so if this sounds good to you, see what we're up to ... You can work in San Francisco, CA; New York, NY; or Seattle, WA You Will: * Apply causal inference ...

Showing results 21-40

What Can You Do With A Phd In Rhetoric information

What are popular job titles related to What Can You Do With A Phd In Rhetoric jobs?

For What Can You Do With A Phd In Rhetoric jobs, the most frequently searched job titles are:

Reinforcement Learning Environment Engineer (Contract)

Hayward, CA โ€ข On-site

Other

Posted 9 days ago


Job description

About the role:

Cobalt is seeking people who can build the environments frontier labs train and evaluate agents in: tasks with real difficulty, unambiguous success conditions, and scoring that survives contact with a capable model.

This opportunity is suited to reinforcement learning researchers, research engineers, simulation and tooling engineers, and people who have built serious benchmarks, competition problems, or training environments. Depth in RL is valuable, but so is the engineering discipline required to make an environment reproducible and hard to game.

You do not need prior experience in data annotation. What matters is that you can take a domain, decide what a meaningful task in it looks like, and build something that measures it correctly.


What you'll do:

Depending on the project, you may:

  • Design and build task environments with programmatic success criteria, including multi-step and tool-using tasks that cannot be solved by a shortcut
  • Specify reward functions and partial-credit schemes, and stress-test them for the ways a capable agent would exploit them rather than solve the task
  • Produce written reasoning traces and reference solutions showing how a competent human works through the tasks you build
  • Evaluate agent trajectories, identifying the specific step at which behavior goes wrong, and classifying failures into a consistent taxonomy
  • Assess whether a scored result reflects genuine task completion, and flag cases where the environment or the metric is measuring the wrong thing

Projects follow their own guidelines, formatting conventions, and quality standards, and you will work with feedback from reviewers and lab research teams.


Required qualifications:

  • Direct experience with reinforcement learning, agent evaluation, simulation, or benchmark and environment construction, whether in research, industry, or substantial open-source work
  • Strong software engineering ability in Python, sufficient to build reproducible environments, harnesses, and automated scoring
  • A PhD in a quantitative discipline, or equivalent depth demonstrated through published work, open-source contributions, or production systems
  • Understanding of reward hacking and specification gaming, and the instinct to look for them in your own designs before someone else does
  • Ability to explain each step of your reasoning and design decisions clearly in writing


Why join Cobalt AI:

  • Advance frontier AI where it counts. Apply your expertise to data that frontier labs cannot obtain any other way, where your reasoning directly shapes how the next generation of models works through technical problems.
  • Grow professionally. Expand your influence through evaluation projects, advisory roles, and research collaborations, while developing a working understanding of how frontier models are trained and assessed.
  • Work with a top-tier network. Collaborate with researchers and engineers from leading institutions and labs on high-impact, flexible work.
  • Set your own schedule. Flexible 10 to 40 hour weeks that fit around your existing work and your life.
  • Competitive pay. Rates vary by project and are determined by a number of factors, including scope, skillset, and experience.