DeepSeek V4.1 Flash: The 552B Model That Beat the 1.6T Flagship, the Retirement That Lasted a Day, and the 510 GB Question

DeepSeek V4.1 Flash: The 552B Model That Beat the 1.6T Flagship, the Retirement That Lasted a Day, and the 510 GB Question

0 View

Publish Date:
12 September, 2026
Category:
Sports
Video License
Standard License
Imported From:
Youtube

By Evan Vega

DeepSeek Releases High-Efficiency AI Model Amid Accessibility Concerns

DeepSeek has launched DeepSeek-V4.1-Flash, a high-performance mixture-of-experts AI model designed to challenge industry leaders in agentic coding and reasoning. While the company touts the model’s efficiency and competitive pricing, independent analysis reveals significant hardware barriers for users attempting to run the “open weights” software locally.

The new model utilizes a Causal Encoder-Decoder architecture with 552 billion total parameters. To optimize compute, the system activates only 8 billion parameters during prefill and 16 billion during decode. DeepSeek reports that the model was pre-trained on 45 trillion tokens and features a million-token context window and native image processing capabilities.

According to DeepSeek’s internal benchmarks, V4.1-Flash outperforms the company’s own 1.6-trillion-parameter flagship, V4-Pro, in several key categories. On the Terminal-Bench 2.1, the Flash model scored 90.6, surpassing both V4-Pro (87.9) and Claude Opus 5.0 (89.1). It also led in DeepSWE v1.1 with a score of 74.2, as well as in CyberGym and AutomationBench.

The company has positioned the model as a cost-effective alternative for developers, pricing off-peak usage at $0.15 per million input tokens and $0.60 per million output tokens.

Despite the “open weights” designation under the MIT license, critics point out that the model’s physical size makes it inaccessible to the average consumer. The checkpoint’s weight index totals approximately 510 gigabytes, a footprint that exceeds the storage and memory capacities of most consumer-grade hardware. This size is attributed to the 552-billion-parameter backbone and a 196-billion-parameter Engram memory system.

Further confusion surrounded the model’s rollout. DeepSeek initially announced that all V4-Pro requests would be routed to the new Flash model starting Sept. 14, but the company later withdrew that statement in its API changelog.

Technical specifications highlight the use of Compressed Sparse Attention 2, which allows layers to share key-value storage. DeepSeek claims this reduces the global cache footprint to 890 bytes per token, roughly one-quarter of the requirements for the previous V4-Flash version.


Read the full investigation →

Related: Frontier Watch

Read the full file →

.