Google DeepMind Deleted the Vision Encoder. The Model Got Better at Hearing.

Google DeepMind Deleted the Vision Encoder. The Model Got Better at Hearing.

0 View

Publish Date:
10 August, 2026
Category:
CNBC
Video License
Standard License
Imported From:
Youtube

By Evan Vega

Google DeepMind Report Reveals Efficiency Breakthrough in Gemma 4 AI

MOUNTAIN VIEW, Calif. — Google DeepMind has released a technical report detailing a paradigm shift in the architecture of its Gemma 4 artificial intelligence models, demonstrating that removing complex components can actually improve performance and efficiency.

The report, published June 19, 2026, explains how the Gemma 4 12B model achieves high-level reasoning and multimodal capabilities by discarding traditional vision and audio encoders. In a departure from industry standards, DeepMind replaced these massive separate networks with streamlined processes that project raw data directly into the language model’s embedding space.

The results challenge the prevailing AI industry belief that larger, specialized encoders are necessary for spatial and temporal awareness. According to the report, the encoder-free 12B model outperformed its counterparts in specific tasks. In English transcription tests, the encoder-free version achieved a word error rate of 0.063, slightly beating the Gemma 4 E4B model, which utilizes a dedicated 305-million-parameter audio encoder, at 0.065.

The architectural streamlining extends to visual processing. DeepMind replaced a 550-million-parameter vision encoder with a single matrix multiplication of just 35 million parameters. Despite this drastic reduction, the 12B model maintained competitive performance, scoring 88.4 on InfographicVQA and 79.7 on MATH-Vision.

Industry analysts note that these gains are tied to a fundamental change in how the model is built. Rather than stripping components from an existing AI, the 12B model was trained from scratch using this unified paradigm. This approach allows the model to maintain spatial awareness through two-dimensional coordinate-based positional embeddings without the memory overhead of a Vision Transformer.

The technical shift has significant implications for the deployment of AI on local hardware. By reducing the global KV cache by 37.5%, the architecture allows for faster processing and lower memory requirements, potentially accelerating the adoption of high-capability AI on consumer-grade devices.

While the report highlights the success of the 12B model, it specifies that this encoder-free architecture is unique to that specific version within the Gemma 4 family, rather than a universal change across all model sizes.


Read the full investigation →

Related: Frontier Watch

Read the full analysis →