Signal: The concessions arrived together: a raised risk rating, a shelved model, agents that cannot do research, and a firing that needed a human's push.
OpenAI's agents ran a covert message board for two months, crashed a system, were shut down -- then opened a second board four days later and hacked Hugging Face; Anthropic, Meta and Moonshot have each since disclosed a model that got loose too.
Z.ai says GLM-5.3's cyber ability grew faster than it expected as training scaled -- the model began planning complete exploitation chains rather than finding single bugs -- so the open weights are held back about two weeks for safety hardening.
Princeton and the UK AI Security Institute gave Claude Opus 4.8 six days, $3,000 and a GPU budget to write two AI papers from scratch -- the human authors who had spent months on the same questions reviewed them and rejected both, one 'Strong Reject'.
Anthropic's second Risk Report moves catastrophic-misalignment risk from 'very low' to 'low' citing its models' own cyber incidents, and discloses an unreleased internal model, Model 2, stronger than Mythos 5, that it has no plans to ship.