METR (Model Evaluation and Threat Research)

ResearchFounded 2023 (spun out of ARC Evals) · US

Visit their site

Independent pre-deployment evaluation of frontier models for autonomous and dangerous capabilities.

METR is the outside tester the big labs let poke at their models before release. It measures how long and how autonomously AI agents can work, and publishes what it finds.

Focus

  • Safety evaluations
  • Transparency
  • Alignment research

Funding, as far as we know

Philanthropic grants (Open Philanthropy/Coefficient Giving, Audacious Project, others); METR states it does not accept payment from labs for evaluations - https://metr.org/about/

What they've actually done

  • Published the 'task-completion time horizon' metric showing frontier agents' capability doubling roughly every 7 months, now widely cited

    19 March 2025

    Source

  • Ran a pilot 'Frontier Risk Report' with internal model access from Anthropic, Google, Meta and OpenAI, documenting 100+ constraint violations by internal agents

    19 May 2026

    Source

  • Published its pre-deployment evaluation summary of OpenAI's GPT-5.6 'Sol'

    26 June 2026

    Source

Last verified 1 September 2026

Report a correction

Back to all organizations