Back
DeepSeekSeptember 22, 20262 sources

DeepSeek Plans 8-Trillion-Parameter Model, Nearly Triple the Largest Open Model

AI Analysis

DeepSeek is signaling a return to raw scale. CEO Liang Wenfeng revealed to investors that the company is currently training a 2-trillion-parameter model and targeting an 8-trillion-parameter flagship — a figure that would make it nearly three times larger than Moonshot AI's Kimi K3, currently the largest open-source model at 2.8 trillion parameters. If realized as an open-weight release, it would dramatically raise the ceiling for the open-model ecosystem.

The ambition sits in tension with DeepSeek's simultaneous efficiency narrative around V4.1-Flash, which emphasized doing more with less memory. Together the two threads suggest a bifurcated strategy: hyper-efficient Flash models for cost-sensitive agentic deployment, and a massive frontier model to compete on raw capability with GPT-6 Astra, Gemini, and Claude. As the company itself noted, total parameter count does not directly equate to active parameters used in every inference — MoE architectures activate only a fraction — so the 8T figure is as much a statement of intent and compute access as a capability guarantee.

The timing matters given DeepSeek's hardware plans: 160,000 Huawei Ascend 950DT chips in Inner Mongolia with training silicon expected in Q4 2026. Training an 8T-parameter model demands enormous compute, and DeepSeek's ability to secure domestic chips amid export controls is the gating factor. For the broader industry, the announcement is a reminder that the scaling race is far from over even as efficiency and agent economics dominate the conversation — and that Chinese labs intend to compete at the very top of the parameter curve, not just on cost.

Sources
AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog