Stories about Qwen3.8-Flash-Next
4 related stories
Experiment with Qwen3.8-Flash-Next 176B Model on NVIDIA GB300 NVL72 for Agentic Coding
AI InsightAlibaba has released the Qwen3.8-Flash-Next model weights as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate. The model uses a multi-modal mixture of experts (MoE) architecture, with a total of 176 billion parameters, including 51 billion N-gram embedding parameters, and 6 billion parameters activated per token. Its native context window size is 262,144 tokens, and can be extended to 1 million tokens with YaRN.Experiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding
AI InsightAlibaba has released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture, allowing developers to experiment and evaluate it. The model adopts a multi-modal MoE architecture, with 12.5 billion main model parameters, supplemented by 5.1 billion N-gram embeddings, and 600 million activation parameters per token. It supports a context window of up to 262,144 tokens, which can be extended to a maximum of 1 million tokens.Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture
AI InsightAlibaba's Qwen team has released the Qwen3.8-Flash-Next model, which features over 12.5 billion multi-modal MoE and 600 million active parameters. The model is a preview of the Qwen4 architecture, incorporating technologies such as Gated DeltaNet and Qwen Sparse Attention hybrid architecture, Gated Residual, N-gram Embedding, and Muon optimizer. The model's training cost is 9 times lower than Qwen3.7-Plus and requires a 172.78 GiB FP8 checkpoint.Alibaba releases Qwen3.8-Flash-Next, targeting "ultimate cost efficiency"
AI InsightAlibaba's Qwen team released Qwen3.8-Flash-Next, a hybrid expert model that only activates 6 out of 125 billion parameters at a time. With a training cost that is only 1/9 of the cost, it surpassed larger competitors like DeepSeek-V4-Flash and Claude Opus 4.6 in encoding and office benchmarks. This achievement puts greater price pressure on companies like OpenAI and Anthropic.