Mistral Small 4 Sets Sparse-Efficiency Benchmark: 119B Total, 6.5B Active Parameters

Mistral Small 4 pushes sparse mixture-of-experts efficiency: 119 billion total parameters but only 6.5 billion active per forward pass, meaning it delivers larger-model quality at a fraction of the compute cost. At FP4 precision it fits in about 142GB of GPU memory, positioning it as an open-source option for teams that want strong capability without frontier-scale inference bills — directly relevant amid the week's DeepSeek- and Qwen-driven price war.
Beyond the model, Mistral rounded out its stack: it acquired Austrian physics-AI group Emmi AI, expanding into scientific/simulation AI, and shipped Voxtral TTS for speech synthesis, an OCR 4 document tool, and a Search Toolkit. The breadth signals Mistral's ambition to be a full-stack European AI provider rather than a pure model lab.
Competitively, the extreme active-parameter sparsity is the headline technical claim — running near-frontier quality on a single high-memory node is exactly what self-hosting European enterprises and privacy-sensitive customers want. It also lands in a crowded open-weight week alongside Qwen 3.8-Max and MiniMax H3, with Simon Willison noting the flurry of near-simultaneous releases.
Caveats: 6.5B active parameters implies strong routing, but MoE models can underperform dense models on tasks that need broad cross-domain reasoning, and Mistral's benchmark claims await independent confirmation. The Emmi AI acquisition is early-stage and its payoff unclear. What to watch: third-party evals of Small 4 against comparable open models, and whether the Voxtral/OCR/Search toolkit gains real developer traction versus incumbents.