阿里巴巴发布Qwen3.8-Flash-Next,瞄准“终极成本效率”
Alibaba releases Qwen3.8-Flash-Next, targeting "ultimate cost efficiency"
阿里巴巴的Qwen团队发布了Qwen3.8-Flash-Next:一个125亿多模态MoE,具有6亿活跃参数,预览Qwen4架构
Alibaba's Qwen team has released the Qwen3.8-Flash-Next model, which features over 12.5 billion multi-modal MoE and 600 million active parameters. The model is a preview of the Qwen4 architecture, incorporating technologies such as Gated DeltaNet and Qwen Sparse Attention hybrid architecture, Gated Residual, N-gram Embedding, and Muon optimizer. The model's training cost is 9 times lower than Qwen3.7-Plus and requires a 172.78 GiB FP8 checkpoint.
Verify source →阿里巴巴发布Qwen3.8-Flash-Next,瞄准“终极成本效率”
Alibaba's Qwen team released Qwen3.8-Flash-Next, a hybrid expert model that only activates 6 out of 125 billion parameters at a time. With a training cost that is only 1/9 of the cost, it surpassed larger competitors like DeepSeek-V4-Flash and Claude Opus 4.6 in encoding and office benchmarks. This achievement puts greater price pressure on companies like OpenAI and Anthropic.
Verify source →