Stories about Qwen Team
1 related stories
Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture
AI InsightAlibaba's Qwen team has released the Qwen3.8-Flash-Next model, which features over 12.5 billion multi-modal MoE and 600 million active parameters. The model is a preview of the Qwen4 architecture, incorporating technologies such as Gated DeltaNet and Qwen Sparse Attention hybrid architecture, Gated Residual, N-gram Embedding, and Muon optimizer. The model's training cost is 9 times lower than Qwen3.7-Plus and requires a 172.78 GiB FP8 checkpoint.