Hugging Face's State of Open Models report: small models still dominate real-world usage

Hugging Face's 'State of Open Models: Summer 2026' report offers a data-grounded snapshot of an ecosystem moving fast in two directions at once. The headline finding: frontier models keep getting larger, but small models still dominate actual real-world usage. Qwen leads local inference, followed by Google's Gemma, and Chinese labs — particularly Alibaba's Qwen — set the pace on downloadable models. AI agents have become a major force on the Hub, reflecting the industry-wide shift toward agentic workloads.
The report quantifies the ecosystem's extreme concentration: from January to August 2026, public model repositories, datasets, and Spaces all grew substantially, yet just 1.5% of repositories account for 99.2% of all downloads. That power-law distribution means the vast majority of the millions of models on the Hub see negligible use, while a tiny handful — the popular Qwen, Gemma, and Llama derivatives — absorb nearly all demand.
CEO Clement Delangue reinforced the local-AI theme separately, noting Transformers.js crossed 10 million monthly downloads (nearly 10x in six months) as the most popular library for running models directly in the browser. Alibaba's Qwen account celebrated its local-inference lead in the report.
The report is a useful corrective to frontier-model hype: for practitioners, a well-quantized 27B or 30B model running locally often beats calling a trillion-parameter API on cost and latency. The skewed download distribution also flags an ecosystem risk — a handful of labs (increasingly Chinese) effectively define what 'open' means in practice. Watch whether Meta's Muse Glimmer and NVIDIA's Nemotron Lightning can crack the top tier that Qwen and Gemma currently own.