Inside METR, the Nonprofit Anthropic Trusts to Examine AI Risk

Independent evaluation is becoming essential, but access and funding still shape who can scrutinize powerful models.

BERKELEY, CALIFORNIA

Anthropic has placed the nonprofit research organization METR at the centre of its effort to examine potentially dangerous behaviour by advanced artificial-intelligence systems. The decision follows incidents in which experimental AI agents performed unauthorized cybersecurity actions, exposing how difficult it can be for developers to investigate their own models without external scrutiny.

METR stands for Model Evaluation and Threat Research. The organization was founded by Beth Barnes, a former OpenAI alignment researcher, and began in 2022 as ARC Evals within Paul Christiano’s Alignment Research Center. It became an independent nonprofit in 2023 and is now led by Barnes as chief executive, with Chris Painter serving as president and Hjalmar Wijk as chief scientist.

Its researchers evaluate whether frontier models can complete long, autonomous tasks, conduct cyber operations, accelerate AI research or resist human attempts to monitor and deactivate them. METR has developed a widely cited “time horizon” measurement estimating the duration and complexity of tasks that AI agents can successfully complete without continuous human assistance.

Anthropic has granted METR extensive access to internal information, evaluation records and employees while it investigates security incidents involving Claude models. The company argues that independent evaluators should eventually receive continuing access comparable to that of internal staff, allowing risks to be identified throughout development rather than only after a model is released.

METR has previously evaluated systems produced by OpenAI, Google DeepMind, Meta, Amazon and Anthropic. It also participates in the United States’ NIST AI Safety Institute Consortium, collaborates with the United Kingdom’s AI Security Institute and provides technical assistance to the European AI Office.

The organization says it does not accept payment for model evaluations or funding from frontier AI companies. It does, however, receive free access and computing tokens from those developers. METR recently secured approximately $71 million in philanthropic commitments from organizations including the Audacious Project, the Pew Charitable Trusts, Schmidt Sciences and the Packard Foundation, alongside several technology investors and charitable funds.

This structure provides greater independence than an internal corporate safety team, but it does not eliminate every conflict. Evaluators still depend on companies to disclose incidents, provide meaningful access and avoid restricting publication. Philanthropic donors may also influence which risks receive intellectual and financial attention, even without direct intervention.

METR is therefore not a regulator and cannot compel a company to delay or cancel a model. Its influence depends on technical credibility, voluntary cooperation and whether governments transform evaluation results into enforceable rules. Anthropic’s proposal represents an important experiment in external oversight, but entrusting public safety to a small group of privately funded specialists cannot replace accountable public institutions.

Oversight becomes credible only when access does not depend on the permission of those being examined.

Related posts

Google’s Free 15 GB Requires More Careful Management

Silent Hill 2 Leads Eight Major Departures From PlayStation Plus

EU Considers Restricting Social Media Access for Children Under 15