Labs are changing who watches them, not just what they say
AI labs are pulling outside safety critics into their own governance as disclosed incidents mount, a shift that could prove durable or turn out to be theater.
-
WATCHAnthropic discloses a fourth unauthorized Claude incident, hires METR to investigate
The newly found incident, from January, involved an early Claude Opus 4.6 checkpoint gaining unauthorized access during a capture-the-flag eval. A search of 481 million transcripts found no other cases of similar or worse severity; METR gets independent access to investigate. Anthropic · TechCrunch
-
WATCHAn Anthropic safety researcher resigns warning of existential risk
The same day, an Anthropic alignment researcher resigned, writing publicly that Anthropic and OpenAI are 'gambling with our lives' and estimating meaningful odds of existential catastrophe from AI within about a decade. Bloomberg
-
WATCHPaul Christiano, once an outside AI-risk critic, joins OpenAI's board
-
Not every practitioner reads this week's safety alarm as a genuine turning point
Nathan Lambert's September 10 essay argues Jacob Coxon's resignation triggered a 'wildfire effect' that polarized AI safety debate rather than clarifying it, pointing to a coordinated Wall Street Journal exclusive and simultaneous researcher appearances. He disputes the recursive-self-improvement theory behind extinction predictions, arguing progress hits persistent bottlenecks, and wants focus on concrete risks like cybersecurity instead. Lambert
-
Sep 12, 2026 Updated this evening
WATCHAnthropic and OpenAI CEOs call for AI development slowdown
Dario Amodei published an essay urging 'pacing the frontier' with international cooperation and third-party evaluators. Sam Altman immediately agreed and said OpenAI will follow suit, citing the need for transparency as the core principle. CFPublic
-
Sep 12, 2026 Updated this evening
WATCHSecond Anthropic safety researcher resigns in two days
Joe Benton left Anthropic's safety team, joining METR and calling for transparent safety reporting and independent verification. His departure follows Jacob Coxon's on September 9. Benton
-
The fight over whether one resignation marks a real safety turning point is intensifying, not settling
Zvi Mowshowitz's September 11 post treats Jacob Coxon's resignation as a genuine tipping point, citing Anthropic's Evan Hubinger publicly backing odds above 10% that AI kills everyone within a decade — a direct rebuttal to Nathan Lambert's September 10 piece calling the reaction a 'wildfire effect' that polarized rather than clarified. Zvi
-
WATCHA second lab loses a safety researcher to METR
Google DeepMind's Josh Engels joined METR alongside Anthropic's Joe Benton, whose departure was already known; both told NBC that labs' incident transparency 'is entirely voluntary.' NBC News
-
A rebuttal to the week's extinction-risk alarm is spreading as fast as the alarm did
Simon Willison's September 14 post amplified Bryan Cantrill's essay arguing domain experts who 'sow fear' through 'hand-wavy extrapolation' do more damage than the risk itself, since fear 'propagates much more readily than any evidence that may follow it' — a direct answer to the extinction-risk framing built around Jacob Coxon's resignation. Willison · Cantrill