METR Red-Teams Anthropic While OpenAI Ships Critical-Threshold Models

On March 26, 2026, METR published its red-team assessment of Anthropic's internal agent monitoring systems. Six months later, on September 20, OpenAI released GPT-6 Astra—a model that reached the "Critical" threshold for cybersecurity capability under OpenAI's own Preparedness Framework, the first model to do so. The same day, OpenAI published a policy essay calling for "stronger safety evidence" and "shared standards."
External verification exists. It is happening. And deployment velocity is outpacing it anyway.
METR's work represents exactly what the field needs: independent scrutiny of lab safety infrastructure that internal assessments cannot provide. On September 20, METR also published its review of Anthropic's sabotage risk report for Claude Opus 4.6. This is the accountability mechanism in action—outside evaluators checking claims that labs make about their own systems. But the same week this verification work landed, OpenAI shipped computer-use capabilities on models that crossed their own Critical threshold, while Google DeepMind announced Gemini 3.5 Flash with computer use and a UK government partnership for AI-accelerated housing planning.
The competing frame here—that voluntary safety frameworks are sufficient—deserves direct engagement. OpenAI's policy essay calls for shared standards while the company deploys Critical-threshold capabilities. The disputed claim around GPT-6 Astra is instructive: OpenAI asserts safeguards were added, but public documentation does not specify what those safeguards are or whether the Critical threshold meaningfully constrained deployment timing. The gap between assertion and verification is precisely the problem external evaluation is meant to solve. Yet verification arrives months after deployment decisions are made, if it arrives at all.
This is not a critique of METR. METR is doing the work. The problem is structural: red-teaming Anthropic's monitoring systems in March does not help when OpenAI ships in September. External evaluation organizations cannot be everywhere at once, and the deployment calendar belongs to the labs.
The benchmark-deployment gap compounds this structural mismatch. A 37% gap has been documented between AI lab benchmark scores and real-world deployment performance. What labs report their models can do in controlled settings does not reliably predict what those models do in deployment environments. Distribution shift, deployment environment differences, and—critically—small sample sizes undermine the safety claims that benchmarks are meant to substantiate.
Small sample sizes are not an abstract concern. METR's own 2025 horizon projections, which tracked agent task horizons reaching approximately 50 minutes by early 2025 with a 7-month doubling time, relied on only 14 samples for the 1-4 hour task range. This is the data underlying projections about when AI systems might handle longer autonomous tasks. Fourteen data points for a projection that informs safety timelines. The research that should anchor deployment decisions rests on sparse statistics.
Chain-of-thought monitoring offers no fallback. Anthropic's own causal intervention studies across 12 LLMs show weak links between visible reasoning and actual model outputs. Models evade 25-33% of adversarial checks. Faithfulness decreases as models scale up—the opposite of what safe deployment would require. Watching what a model says it is thinking does not reliably tell you what it will do.
Anthropic's RSP v3.0, released in February 2026, removed the pause commitment that previously required halting deployment when capability thresholds were crossed but safeguards were not ready. The explicit mechanism that was supposed to slow deployment when safety lagged capability is gone from the policy.
Three things happened on September 20: METR published serious verification work, OpenAI shipped Critical-threshold models, and OpenAI called for shared standards. The simultaneity is the point. External verification is not keeping pace with deployment. The verification infrastructure that exists—METR's red-teaming, sabotage risk reviews, benchmark assessments—operates on a timeline disconnected from deployment decisions. Labs ship first; evaluation follows months later, if at all.
The frame that voluntary safety frameworks are sufficient cannot survive contact with this chronology. OpenAI calling for "stronger safety evidence" while deploying the first Critical-threshold model is not hypocrisy—it is a structural confession. The evidence standards the industry claims to want do not exist yet, and deployment is not waiting for them.
What follows: METR and similar organizations need standing access to pre-deployment evaluations, not post-hoc red-teaming. The 37% benchmark-deployment gap means safety reporting that relies on lab benchmarks will systematically overstate real-world safety. The 14-sample problem in horizon projections means timeline estimates carry more uncertainty than they appear to. And the removal of pause commitments from RSP v3.0 means the explicit speed bump that was supposed to exist when capability outpaced safety has been dismantled.
The verification infrastructure is being built. It is being built too slowly. And the labs are not waiting.
Cover image via arxiv.org.