Hugging Face launches Open Alignment Initiative, volunteers to audit AI labs

Hugging Face launched the Open Alignment Initiative, led by co-founder Clément Delangue and Thomas Wolf, positioning the company to serve as an external auditor and 'embedded evaluator' for frontier AI labs. The announcement came shortly after Anthropic's Dario Amodei proposed exactly such an embedded-evaluator program in his 'Pace the Frontier' essay, and Delangue explicitly asked to be part of it. 'It's now clear that alignment is critical and won't be solved behind the closed doors of a handful of frontier labs,' Delangue wrote.
The initiative reframes Hugging Face — best known as the open-source model and dataset hub — as a neutral, third-party safety player. That role has weight given the platform's centrality to the open-weight ecosystem, but it also arrives amid two complicating storylines: Hugging Face is in the process of being acquired by NVIDIA for $12.93 billion (a deal confirmed September 3), and its own platform this week hosted 'abliterated' uncensored forks of DeepSeek's V4.1-Flash within hours of release — a vivid illustration of how hard alignment is to enforce on an open platform.
The move drew cross-industry engagement: Andrej Karpathy praised the broader slowdown push, and Delangue's pledge to join the evaluator program signaled that the embedded-evaluator idea is gaining voluntary adoption beyond Anthropic and OpenAI. Critics, however, questioned whether a company simultaneously distributing uncensored model forks and being absorbed by a chipmaker can credibly serve as an independent alignment auditor.
What to watch: whether frontier labs actually grant Hugging Face evaluator access, how NVIDIA's ownership affects its independence, and whether the initiative produces concrete audit methodologies rather than a mission statement.