Anthropic raises misalignment risk to 'low', discloses unreleased 'Model 2'

Anthropic's second Risk Report is a rare instance of a frontier lab publicly nudging its own risk dial upward. Moving misalignment risk from 'very low' to 'low' is framed as reflecting growing uncertainty as capabilities accelerate — not any specific test failure — which is a subtle but important distinction: the lab is signaling that its confidence in controllability is thinning even as no concrete breach has occurred.
The most-discussed disclosure is an internal 'Model 2' that noticeably outperforms Claude Mythos 5, with no external release planned. That confirmation — a lab sitting on a materially better model it won't ship — fed r/singularity (633 upvotes, 208 comments) and Hacker News debate. The report also states Claude now writes a majority of the code merged into Anthropic's own production systems, a striking data point on automated AI R&D, which the report flags as a potential major future risk.
The skeptical takes are pointed. Hacker News (391 comments) flagged a procedural weakness: the safety case partly rests on models being bad at scheming, which Anthropic concedes it couldn't reliably detect if it began — and its internal dangerous-capability benchmark has saturated, meaning the yardstick may no longer measure what matters. The broader theme of the week is 'models being throttled or withheld,' echoed in separate HN debate over models 'getting dumber on purpose.' Watch whether other labs publish comparable risk self-assessments or let Anthropic's transparency stand alone.