Qwen releases Qwen3.8-Flash-Next as an early preview of Qwen4's architecture
2026-08-26
Simon Willison reported Qwen released Qwen3.8-Flash-Next, described as a multimodal MoE model serving as 'an early preview of the architecture used in Qwen4,' with 125B total parameters and 6B active parameters. Willison tested Unsloth-quantized versions (72.5GB and 78.9GB) on an Nvidia DGX Spark.
Significance 3: A real open-weights release previewing a tracked lab's next-generation architecture is worth reading; held to 3 because the only source is a single curator's informal hands-on testing with no benchmark numbers.
Implications · machine-drafted, not owner judgment
Releasing an early Qwen4-architecture preview as open weights, rather than holding the new architecture back for a closed flagship, continues Qwen's pattern of open-sourcing at the frontier of its own roadmap rather than trailing it — useful signal for the open-weights-gap thesis even before benchmarks exist.
- Formal Qwen4 flagship release and how much it diverges from this preview architecture
- First independent benchmark numbers for Qwen3.8-Flash-Next
Sources
- curator Qwen3.8-Flash-Next retrieved 2026-08-27