A 153,000-View DeepSeek V4.1 Video, Checked: Mostly Right, One Wrong Row, and 510 GB Nobody Mentioned

A 153,000-View DeepSeek V4.1 Video, Checked: Mostly Right, One Wrong Row, and 510 GB Nobody Mentioned

0 View

Publish Date:
13 September, 2026
Category:
Fox News
Video License
Standard License
Imported From:
Youtube

By Evan Vega

Audit Finds Most Viral DeepSeek V4.1 Analysis Accurate Despite Key Omissions

A detailed audit of the most-watched independent explanation of DeepSeek’s new V4.1 model has found the analysis to be largely factual, though it contained critical errors regarding competitive performance benchmarks.

The video, titled “Deepseek did it again…”, was published Sept. 11 to a channel with 635,000 subscribers. By the following morning, the content had amassed 153,218 views, establishing it as the primary source of public information on the model’s release. An audit conducted by Frontier Watch verified the claims against DeepSeek’s official README, weight index, and pricing documentation.

The audit confirmed that the majority of the technical specifications were reported accurately. This includes the model’s 552B mixture-of-experts architecture, featuring 8 billion active parameters for input and 16 billion for output. The analysis correctly cited a DeepSWE v1.1 score of 74.2, surpassing Claude Opus 5.0 (74.0) and GPT-5.6 Sol (73.0), as well as a leading CyberGym score of 88.1.

Financial and hardware claims also held up to scrutiny. The report accurately quoted pricing at $0.15 per million input tokens off-peak and $0.60 for output, while correctly noting the model’s reduced hardware requirements, including a global KV cache of 890 bytes per token.

However, the audit identified a significant factual error regarding the “Terminal-Bench 3.0” benchmark. The video claimed only Claude Opus 5.0 outperformed DeepSeek; in reality, both Opus 5.0 (43.3) and GPT-5.6 Sol (34.4) beat DeepSeek’s score of 30.0. While DeepSeek leads in the older Terminal-Bench 2.1 version, it ranks third in the more recent 3.0 and 4.0 versions, trailing by 13.3 and 20.6 points, respectively.

The audit also highlighted a misleading statement regarding hardware requirements, warning that users acting on the video’s “enough VRAM” claims may find the information insufficient for actual implementation.

Despite these errors, the audit noted that the video provided the only independent evaluation of the model’s actual performance available as of Sept. 12. The presenter specifically highlighted the model’s struggles with “ExploitGym,” where DeepSeek scored 15.3—significantly lower than GPT-5.6 Sol’s 33.7 and Opus 5.0’s 22.1—providing a level of transparency absent from other early coverage.


Read the full investigation →

Related: Frontier Watch

Read the full audit →

.