DeepSeek V4.1 Flash shrinks US lead over Chinese AI to 3%, Bloomberg Intelligence says

Bloomberg Intelligence's latest benchmark tracking puts the gap between the best US and Chinese frontier models at roughly 3 percentage points. That is a sharp compression from about 9% in May and about 15% at the start of 2026. The proximate cause is DeepSeek V4.1 Flash, a multimodal model released September 10. It sits sixth on LiveBench and scores 74.2% on DeepSWE v1.1 and 90.9 on GPQA Diamond.
V4.1 Flash is positioned as a throughput-optimized model at aggressive token prices, and it ships as open weights on Hugging Face under the permissive MIT license. That combination matters as much as the score. Enterprises can self-host a model within a few points of the US frontier with no licensing friction, which removes much of the premium US labs charge for closed API access.
The competitive picture is crowded. On DeepSWE v1.1, V4.1 Flash beats xAI's Grok 4.7 (71.0%) but trails Google's restricted Gemini 4 Argon (77.9%). Comparison posts also note that Argon has no public release date, while DeepSeek's weights are downloadable today. That contrast feeds the week's debate over gated versus open frontier releases. Alibaba's simultaneous announcement that Qwen 4 is in training adds to the sense that Chinese labs are closing in on multiple fronts.
There are caveats. Benchmark parity is not deployment parity, and 'benchmaxxing' accusations are flying at labs on both sides of the Pacific this week. Concerns about Chinese open models persist: a CBS-reported study found pro-Beijing bias in Qwen, and Ethan Mollick noted wryly that even Europe's new 'sovereign' model was fine-tuned on GLM- and Qwen-generated data. The next things to watch are whether US labs respond on price, and whether export-control pressure, already visible in RTX 5090 purchase paperwork, tightens further.