About METR We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation ...
60 Metr Jobs Hiring Near You
About METR We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation ...
About METR We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation ...
About METR We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation ...
Security Engineer
Berkeley, CA · On-site
$285.55 - $503.12/hr
About METR We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation ...
Security Engineer
Berkeley, CA · On-site
$285.55 - $503.12/hr
About METR We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation ...
About METR We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation ...
About METR We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation ...
If you don't fit those categories but believe you can potentially help METR with its work, please also feel free to use this to register your interest. Here is a list of potential future roles.
If you don't fit those categories but believe you can potentially help METR with its work, please also feel free to use this to register your interest. Here is a list of potential future roles.
AISST/MAIA Workshop Lead
Cambridge, MA · On-site
$100K - $150K/yr
The AISST/MAIA Spring 2026 workshops, for example, featured speakers from Anthropic, METR, Redwood Research, Google DeepMind, the AI Futures Project, IFP, CNAS, Eleos AI, RAND, Forethought, CAIS ...
New
AISST/MAIA Workshop Lead
Cambridge, MA · On-site
$100K - $150K/yr
The AISST/MAIA Spring 2026 workshops, for example, featured speakers from Anthropic, METR, Redwood Research, Google DeepMind, the AI Futures Project, IFP, CNAS, Eleos AI, RAND, Forethought, CAIS ...
New
Director of Safety
Brampton, ON · On-site
Monitor, track, and analyze safety metr Accommodations are available upon request for all individuals with disabilities taking part in the recruitment and selection process.
Quick apply
Director of Safety
Brampton, ON · On-site
Monitor, track, and analyze safety metr Accommodations are available upon request for all individuals with disabilities taking part in the recruitment and selection process.
Wildcard
San Francisco, CA · On-site
Our team comes from Anthropic, McKinsey, METR, Weights & Biases, Perplexity, Thiel Fellowship, etc. We've raised $55m from Ribbit, First Harmonic, and Nat Friedman. Read more here, here, and here.
Wildcard
San Francisco, CA · On-site
Our team comes from Anthropic, McKinsey, METR, Weights & Biases, Perplexity, Thiel Fellowship, etc. We've raised $55m from Ribbit, First Harmonic, and Nat Friedman. Read more here, here, and here.
Member of Technical Staff
San Francisco, CA · On-site
$200K - $400K/yr
Our team comes from Anthropic, McKinsey, METR, Palantir, Amazon - 50% are previous founders. We've raised $55m from Ribbit, First Harmonic, and Nat Friedman. Read more here, here, and here. The role ...
Member of Technical Staff
San Francisco, CA · On-site
$200K - $400K/yr
Our team comes from Anthropic, McKinsey, METR, Palantir, Amazon - 50% are previous founders. We've raised $55m from Ribbit, First Harmonic, and Nat Friedman. Read more here, here, and here. The role ...
Software Engineer, RL Environments
$180K - $220K/yr
A background in AI safety or benchmarking organizations (e.g., METR, Artificial Analysis) * A genuine obsession with how data structure, selection, and quality drive model behavior * An ability to ...
Quick apply
Software Engineer, RL Environments
$180K - $220K/yr
A background in AI safety or benchmarking organizations (e.g., METR, Artificial Analysis) * A genuine obsession with how data structure, selection, and quality drive model behavior * An ability to ...
Delivery
San Francisco, CA · On-site
Our team comes from Anthropic, McKinsey, METR, Weights & Biases, Perplexity, Thiel Fellowship, etc. We've raised $55m from Ribbit, First Harmonic, and Nat Friedman. Read more here, here, and here.
Delivery
San Francisco, CA · On-site
Our team comes from Anthropic, McKinsey, METR, Weights & Biases, Perplexity, Thiel Fellowship, etc. We've raised $55m from Ribbit, First Harmonic, and Nat Friedman. Read more here, here, and here.
Operations Lead
San Francisco, CA · On-site
$120K - $300K/yr
Our team comes from Anthropic, McKinsey, METR, Palantir, Amazon - 50% are previous founders. We've raised $55m from Ribbit, First Harmonic, and Nat Friedman. Read more here, here, and here. The Role ...
Operations Lead
San Francisco, CA · On-site
$120K - $300K/yr
Our team comes from Anthropic, McKinsey, METR, Palantir, Amazon - 50% are previous founders. We've raised $55m from Ribbit, First Harmonic, and Nat Friedman. Read more here, here, and here. The Role ...
Software Engineer - RL Environments
San Francisco, CA · On-site
$180K - $220K/yr
Major plus if they've worked for/interned for any RL environment companies in the past or any AI safety or benchmarking orgs like METR, Artificial Analysis, etc.. * Genuine obsession with how data ...
Software Engineer - RL Environments
San Francisco, CA · On-site
$180K - $220K/yr
Major plus if they've worked for/interned for any RL environment companies in the past or any AI safety or benchmarking orgs like METR, Artificial Analysis, etc.. * Genuine obsession with how data ...
Executive Assistant
San Francisco, CA · On-site
Our team comes from Anthropic, McKinsey, METR, Weights & Biases, Perplexity, Thiel Fellowship, etc. We've raised $55m from Ribbit, First Harmonic, and Nat Friedman. Read more here, here, and here.
Executive Assistant
San Francisco, CA · On-site
Our team comes from Anthropic, McKinsey, METR, Weights & Biases, Perplexity, Thiel Fellowship, etc. We've raised $55m from Ribbit, First Harmonic, and Nat Friedman. Read more here, here, and here.
Talent Manager - Full-Time Finance & Accounting Engagement Professionals
New York, NY · On-site
$64K - $79K/yr
JOB REQUISITION Talent Manager - Full-Time Finance & Accounting Engagement Professionals LOCATION NY MIDTOWN NEW YORK Robert Half is looking for professionals to join our Full-Time Engagement ...
Talent Manager - Full-Time Finance & Accounting Engagement Professionals
New York, NY · On-site
$64K - $79K/yr
JOB REQUISITION Talent Manager - Full-Time Finance & Accounting Engagement Professionals LOCATION NY MIDTOWN NEW YORK Robert Half is looking for professionals to join our Full-Time Engagement ...
Environmental Project Manager
Arlington, VA · On-site
Overview What You'll Be Doing As Environmental Project Manager, you will support Cadmus' rapidly growing portfolio of task orders under federal and state contracts, including work with the EPA and ...
Environmental Project Manager
Arlington, VA · On-site
Overview What You'll Be Doing As Environmental Project Manager, you will support Cadmus' rapidly growing portfolio of task orders under federal and state contracts, including work with the EPA and ...
Founder's Office
San Francisco, CA · On-site
Our team comes from Anthropic, McKinsey, METR, Palantir, Amazon - 50% are previous founders. We've raised $55m from Ribbit, First Harmonic, and Nat Friedman. Read more here, here, and here. The Role ...
New
Founder's Office
San Francisco, CA · On-site
Our team comes from Anthropic, McKinsey, METR, Palantir, Amazon - 50% are previous founders. We've raised $55m from Ribbit, First Harmonic, and Nat Friedman. Read more here, here, and here. The Role ...
New
Founding Talent Partner
San Francisco, CA · On-site
$120K - $250K/yr
Our team comes from Anthropic, McKinsey, METR, Weights & Biases, Perplexity, Thiel Fellowship, etc. We've raised $55m from Ribbit, First Harmonic, and Nat Friedman. Read more here, here, and here.
Founding Talent Partner
San Francisco, CA · On-site
$120K - $250K/yr
Our team comes from Anthropic, McKinsey, METR, Weights & Biases, Perplexity, Thiel Fellowship, etc. We've raised $55m from Ribbit, First Harmonic, and Nat Friedman. Read more here, here, and here.
Enterprise GTM
San Francisco, CA · On-site
$300K - $400K/yr
Our team comes from Anthropic, McKinsey, METR, Weights & Biases, Perplexity, Thiel Fellowship, etc. We've raised $55m from Ribbit, First Harmonic, and Nat Friedman. Read more here, here, and here.
Enterprise GTM
San Francisco, CA · On-site
$300K - $400K/yr
Our team comes from Anthropic, McKinsey, METR, Weights & Biases, Perplexity, Thiel Fellowship, etc. We've raised $55m from Ribbit, First Harmonic, and Nat Friedman. Read more here, here, and here.
Commodity Manager
Highland Heights, KY · On-site
Prysmian is the world leader in the energy and telecom cable systems industry. Each year, the company manufactures thousands of miles of underground and submarine cables and systems for power ...
Commodity Manager
Highland Heights, KY · On-site
Prysmian is the world leader in the energy and telecom cable systems industry. Each year, the company manufactures thousands of miles of underground and submarine cables and systems for power ...
Full-time
PTO
Re-posted yesterday
Job description
We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation and misalignment.
METR has consistently set precedents for catastrophic AI risk evaluations, including the first independent safety evaluations (working informally with Anthropic and OpenAI in 2022), the first loss-of-control evaluations and first agentic dangerous capability evaluations, the first evaluations using finetuning (mentioned briefly here), the first independent evaluations using internal information about training, the first review partnership for company risk analysis, the first embedded redteaming, and the first evaluations of internal deployments.
We've been consulted and/or favorably referenced by groups on opposite ends of various spectra, including a16z, Khosla, Gary Marcus, Obama, and Dean Ball, and are known for producing one of the most positive results on AI capabilities (the time horizon trend) and the most negative (our downlift study). We're generally referenced as the canonical third party assessor, e.g. as the obvious candidate to verify conditional pause agreements, and are trusted with AI incident investigations by frontier labs and governments.Â
We believe it is robustly good for policymakers and civil society to have a clear understanding of risks from AI systems, and we are extremely excited to build a team of ambitious, excellent people to tackle one of the most important challenges of our time.
Security at METR is becoming its own dedicated team, and you would be one of its first hires. It is extremely important that we continue to be an organization that frontier AI labs, governments, and the public trust with sensitive model access and confidential information. As misalignment incidents become more extreme and confidential information about models and frontier AI labs becomes more valuable, we expect to be under increasingly heavy pressure.
For us, security encompasses managing endpoints and securing development environments, cloud platform security, safely sandboxing agents and evaluations, VPN and VPC networking, application code reviews, account provisioning and access control, and helping ensure we use the best practices across all of our workflows.
Offensive security: You would be the first person on the team with an offensive security background. You'll run targeted red-team exercises against our own systems and build automated AI red teaming.
High-context detection and response: You will build AI systems that can quickly triage and respond to threats, both from internal agents and external attackers.
Blue-team engineering: Detection engineering, telemetry pipelines, incident response, and hardening across our cloud infrastructure, endpoints, and identity systems.
Securing a unique attack surface: METR's evaluation infrastructure runs frontier AI agents. In the past, we've run pre-deployment model evaluations - executing untrusted, model-generated code at scale on multi-day tasks.
Enabling bleeding-edge research: You'll work closely with our researchers to make dangerous-capability experiments safe to run. We often face extreme reward hacking and evaluation awareness during our pre-deployment evaluations, and expect internal threats from agents to become more extreme.
METR handles some of the most sensitive artifacts in AI - pre-release frontier model access, confidential lab information, and transcripts with raw chain-of-thought. Labs and policymakers trust us with this because of our security posture, and keeping that trust is necessary for everything else we do.
As AI agents are used more aggressively by malicious actors for cyber offense operations, and METR's salience rises in the public eye, we expect to face increasingly sophisticated attacks. Strengthening security at METR can be one of the highest-leverage roles to ensure third parties continue to have access to confidential information necessary to inform the world about current risks.
METR is one of the first organizations to see and closely study misalignment incidents that involve models breaking out of sandboxes, attacking our infrastructure, manipulating graders, and more. We also may pursue incident investigations embedded in frontier labs, in which case internal experience with similar failures will be critical.
- Deep security expertise: You have strong fundamentals across systems, networks, cloud, and identity.
- Offensive security: You have experience acting like an attacker, whether through red teaming, penetration testing, or adversarial research.
- AI/LLM engineering: You build with AI: agent pipelines, LLM-powered tooling, automated workflows, and understand current limitations of those tools.
- AWS: You should know AWS very well, including a deep understanding of IAM policies.
We don't screen on certifications, degrees, or years of experience.
Detection engineering at scale: Experience with SIEM/detection pipelines, writing and tuning detections, and threat hunting.
Cloud and container security: AWS (especially non-trivial IAM), Kubernetes, and infrastructure-as-code environments.
Incident response: You've led or worked severe incidents, ideally those involving AI agents.
AI security research: Familiarity with prompt injection, agent containment, model supply-chain risks, or red teaming AI systems themselves.
- AWS: cloud-native software platforms
- EKS
- Lambda
- ECS
- IAM (in-depth)
- SQS
- CloudWatch
- SecurityHub & GuardDuty
- PostgreSQL: RLS, serverless Aurora
- Pulumi: IaC
- DataDog: SIEM
- Okta: IdP
- Google Workspace: IdP
- Tailscale: networking
- CrowdStrike Falcon: endpoint security
- The office: Catered lunch and dinner daily; in-office gym and shower
- Relocation support: Stipend for moving to the Bay Area
- Time-off and leave: Unlimited PTO and 21-week parental leave for new parents
- Commuter benefit: Monthly transit/parking stipend and an annual Uber budget
- Professional development benefit: for training, courses, conferences, and AI safety education
- Mental health benefit: for therapy, medication, and other mental health expenses
- Wellness benefit: for gym memberships and other wellness expenses
- Work equipment benefit: for home office and workstation equipment expenses