By Evan Vega
In five weeks, Qwen 3.8 27B moved onto a desk, got twice as fast and had its refusals removed. 24 GB: the native vision-language model runs a 262K-token window in laptop memory because 48 of its 64 layers never cache a token. 2.24x: MTPLX’s multi-token prediction on an M5 Max, with the output distribution unchanged. 2.1 points: the MMLU cost of abliteration, published by exactly one of four builds. Every figure is attributed to whoever measured it.
Related: Frontier Watch