1

Content Moderation Agent Jobs (NOW HIRING)

As we scale our agentic coding and AI agent products, you will ensure these rollouts are designed ... Provide high-level oversight for platform abuse and content moderation (in partnership with growth ...

Software Engineer

Folsom, CA ยท On-site +1

$75.50 - $102/hr

Implement robust AI guardrails, including input/output filters, sensitive data scrubbing, content moderation, and fine-grained, per-agent permission boundaries, ensuring secure and compliant AI ...

Contribute to the design and implementation of multi-agent systems, agentic call flows, and LLM ... mitigation, content moderation, adversarial input handling, and model output validation ...

... Implement guardrails, content moderation, authentication, and security best practices for ... โ€ข Knowledge of AI agent frameworks, function calling, tool integration, and workflow ...

Principal Software Engineer

Chicago, IL ยท On-site

$139K - $186K/yr

... agent orchestration, retrieval augmented generation, and vector databases. * Hands-on experience ... Knowledge of AI safety, responsible AI practices, and content moderation systems * Motivation to be ...

Principal Software Engineer

Charleston, WV ยท On-site +1

$124K - $167K/yr

... agent orchestration, retrieval augmented generation, and vector databases. * Hands-on experience ... Knowledge of AI safety, responsible AI practices, and content moderation systems * Motivation to be ...

Principal Software Engineer

Jackson, MS ยท On-site

$117K - $157K/yr

... agent orchestration, retrieval augmented generation, and vector databases. * Hands-on experience ... Knowledge of AI safety, responsible AI practices, and content moderation systems * Motivation to be ...

New

AI Engineer (DFW Area)

Richardson, TX ยท On-site

$107K - $182K/yr

You will work with state-of-the-art foundation models, RAG architectures, and multi-agent systems ... Integrate safety/guardrail layers (e.g., content moderation APIs, Guardrails AI, Rebuff, custom ...

You will work with state-of-the-art foundation models, RAG architectures, and multi-agent systems ... Integrate safety/guardrail layers (e.g., content moderation APIs, Guardrails AI, Rebuff, custom ...

Lead AI Engineer

Columbia, MD ยท On-site

$125K - $180K/yr

Develop and integrate guardrails (e.g., prompt-injection protections, content moderation, output validation), and safeguards for agent loops (e.g., loop prevention, tool call limits, state validation)

AI Eng Lead/Manager

Columbia, MD ยท On-site

$125K - $180K/yr

Develop and integrate guardrails (e.g., prompt-injection protections, content moderation, output validation), and safeguards for agent loops (e.g., loop prevention, tool call limits, state validation)

Lead AI Engineer

Columbia, MD ยท On-site

$125K - $180K/yr

Develop and integrate guardrails (e.g., prompt-injection protections, content moderation, output validation), and safeguards for agent loops (e.g., loop prevention, tool call limits, state validation)

AI Engineer (DFW Area)

Richardson, TX ยท On-site

$107K - $182K/yr

You will work with state-of-the-art foundation models, RAG architectures, and multi-agent systems ... Integrate safety/guardrail layers (e.g., content moderation APIs, Guardrails AI, Rebuff, custom ...

AI Eng Lead/Manager

Columbia, MD ยท On-site

$125K - $180K/yr

Develop and integrate guardrails (e.g., prompt-injection protections, content moderation, output validation), and safeguards for agent loops (e.g., loop prevention, tool call limits, state validation)

Showing results 41-60

Content Moderation Agent information

See salary details

$31.5K

$59.2K

$95K

How much do content moderation agent jobs pay per year?

As of Sep 13, 2026, the average yearly pay for content moderation agent in the United States is $59,196.00, according to ZipRecruiter salary data. Most workers in this role earn between $43,000.00 and $70,500.00 per year, depending on experience, location, and employer.

What are popular job titles related to Content Moderation Agent jobs?

For Content Moderation Agent jobs, the most frequently searched job titles are:

Staff+ Software Engineer, Safeguards Evals

San Francisco, CA โ€ข On-site

Menlo Ventures
Investment Clubs and Venture Capital Companiesย โ€ขย 11 - 50 employees

$320K - $485K/yr

Other

PTO

Posted 25 days ago


Job description

About the role

How do we know our safety systems actually catch misuse? Anthropic increasingly uses AI to investigate potential misuse of Claude โ€” analyzing real-world traffic to surface bad actors, policy violations, and emerging threats. Its findings inform enforcement actions and model launch decisions, which means we need rigorous, trustworthy answers to questions like: Does the monitoring agent catch what it should? Where does it fail? Does it stay reliable as adversaries adapt, as models improve, and as the agent itself changes?

This role builds the evaluation infrastructure that answers those questions. You'll sit at the intersection of applied ML research and engineering โ€” designing experiments to measure how well an investigative agent performs across harm areas, building datasets that represent real abuse rather than synthetic benchmarks, and shipping those methods into pipelines that gate every change to the system. Your work directly determines how much trust Anthropic can place in its automated abuse detection, and where we invest to make it better.

Key responsibilities
  • Build and own the evaluation harness for an agentic investigation system โ€” defining metrics, test cases and grading approaches for a complex long horizon agent
  • Construct high-quality eval datasets representing real-world misuse across harm areas (e.g., cyber attacks, bio weapons, influence operations), drawing from real traffic patterns and synthetic generation
  • Measure agent performance end-to-end (detection precision/recall, investigation quality, robustness) and drive hill-climbing on the hardest harm areas
  • Analyze coverage to identify measurement gaps, and evolve evals so they remain unsaturated and high-signal as agent capabilities advance
  • Productionize successful research into regression and release pipelines that run on every agent change, prompt update, and underlying model upgrade
  • Build tooling that enables policy experts to author, run, and iterate on evaluations without engineering support
  • Construct RL environments to improve Claudeโ€™s safety investigation capabilities.
Minimum qualifications
  • Proficiency in Python and comfort working across the stack
  • Experience building and maintaining data pipelines
  • Experience working with LLMs and a working understanding of their capabilities and failure modes โ€” especially agentic systems with tool use and multi-step reasoning
  • Strong data analysis skills โ€” you can draw reliable insights from large datasets
  • Ability to move fluidly between research prototyping and production-quality code
  • Ability to translate ambiguous problems into concrete, testable experiments
Preferred qualifications
  • 8+ years of industry software engineering experience
  • Expertise in building or contributing to agent evaluation frameworks, benchmarks, or automated grading systems
  • Extensive experience in trust and safety, content moderation, or abuse detection systems
  • Experience in red teaming, adversarial testing, or jailbreak research on AI systems
  • Experience with synthetic data generation or data augmentation
  • Experience with distributed systems or large-scale data processing
  • Experience with prompt engineering or building LLM-powered applications

The annual compensation range for this role is listed below.

Annual Salary: $320,000 โ€” $485,000 USD

Logistics
  • Minimum education: Bachelorโ€™s degree or an equivalent combination of education, training, and/or experience
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
  • Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
  • Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.
  • Visa sponsorship: We do sponsor visas! However, we arenโ€™t able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.

We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues.

#J-18808-Ljbffr