Stories about Alibaba
6 related stories
Experiment with Qwen3.8-Flash-Next 176B Model on NVIDIA GB300 NVL72 for Agentic Coding
AI InsightAlibaba has released the Qwen3.8-Flash-Next model weights as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate. The model uses a multi-modal mixture of experts (MoE) architecture, with a total of 176 billion parameters, including 51 billion N-gram embedding parameters, and 6 billion parameters activated per token. Its native context window size is 262,144 tokens, and can be extended to 1 million tokens with YaRN.Experiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding
AI InsightAlibaba has released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture, allowing developers to experiment and evaluate it. The model adopts a multi-modal MoE architecture, with 12.5 billion main model parameters, supplemented by 5.1 billion N-gram embeddings, and 600 million activation parameters per token. It supports a context window of up to 262,144 tokens, which can be extended to a maximum of 1 million tokens.Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture
AI InsightAlibaba's Qwen team has released the Qwen3.8-Flash-Next model, which features over 12.5 billion multi-modal MoE and 600 million active parameters. The model is a preview of the Qwen4 architecture, incorporating technologies such as Gated DeltaNet and Qwen Sparse Attention hybrid architecture, Gated Residual, N-gram Embedding, and Muon optimizer. The model's training cost is 9 times lower than Qwen3.7-Plus and requires a 172.78 GiB FP8 checkpoint.Alibaba releases Qwen3.8-Flash-Next, targeting "ultimate cost efficiency"
AI InsightAlibaba's Qwen team released Qwen3.8-Flash-Next, a hybrid expert model that only activates 6 out of 125 billion parameters at a time. With a training cost that is only 1/9 of the cost, it surpassed larger competitors like DeepSeek-V4-Flash and Claude Opus 4.6 in encoding and office benchmarks. This achievement puts greater price pressure on companies like OpenAI and Anthropic.Alibaba's Wan3.0 generates AI videos up to 30 seconds long from text, images, and documents
Thomson Reuters bets $40M on owning its AI instead of renting from OpenAI or Anthropic
AI InsightThomson Reuters invested $40 million to develop its own language model called Thomson, based on Alibaba's Qwen, which performed well on company content, such as Westlaw.行业影响有限