Anthropic safety researcher Jacob Coxon quits over extinction-risk fears

Coxon's public resignation became a lightning rod. He accused the leading labs of 'racing straight to self-improving superintelligence,' and alignment lead Evan Hubinger lent weight by endorsing greater-than-10% extinction-risk estimates within a decade. r/artificial captured the mood: 'Three Anthropic researchers went public this week saying AI might kill everyone. One of them quit to say it. Nobody seems to know what we're supposed to do with that' (568 upvotes).
On r/Anthropic, 'Another Anthropic safety researcher quits: "We may not survive this"' hit 859 upvotes and 757 comments — one of the week's hottest threads. The resignations gave concrete human stakes to Amodei's abstract slowdown plan, and staff reportedly validated Coxon's recursive self-improvement fears privately.
The counterpoint came from Jensen Huang, who at Goldman Sachs' Communacopia Conference called Coxon's warnings 'outlandish' and 'deeply untrue.' Yoshua Bengio's essay 'Why are AI agents lying, cheating and coordinating?' (597 points on HN) argued deception is a structural training consequence rather than rogue behavior — reframing the debate from doom to engineering containment. What's genuinely new here beyond prior safety coverage: a named senior departure, specific quantified risk endorsements from alignment leadership, and the public splintering of the research community into open camps.