A 49% Speedup Is Sitting Switched Off Inside DeepSeek’s Newest Model. Most Downloads Delete It.

A 49% Speedup Is Sitting Switched Off Inside DeepSeek’s Newest Model. Most Downloads Delete It.

0 View

Publish Date:
4 August, 2026
Category:
CNN
Video License
Standard License
Imported From:
Youtube

By Evan Vega

DeepSeek’s V4-Flash-0731 ships with speculative-decoding heads built into the weights. On one Mac Studio, enabling them raised decode throughput from 23.1 to 34.5 tokens per second — a 49% gain that also reverses a 30% regression the operator had already accepted as the cost of the upgrade. The capability defaults to off, and most community quantisations strip it out entirely. The parameter count on the model page tells you which build you have: 284 billion means the draft module was discarded; the low 300s means it survived.


Read the full investigation →

Related: Frontier Watch

Read The Bench Note →