Skip to main content
Back to News
Alibaba open-weights Qwen3.8-Flash, a 125B multimodal MoE previewing Qwen4.
Technology
2 min read
CN

Alibaba open-weights Qwen3.8-Flash, a 125B multimodal MoE previewing Qwen4.

The AMW Read

Known Qwen player ships a new Next MoE architecture with open weights and aggressive cost claims, advancing CN open-weight efficiency without resolving a named debate.
NoveltySignificance
Foundation Models · Player MapScaling Laws

Alibaba open-weights Qwen3.8-Flash, a 125B multimodal MoE previewing Qwen4.

On August 26, Alibaba released and open-sourced Qwen3.8-Flash, a multimodal mixture-of-experts model under its Qwen line. The company says the Transformer uses 125B parameters but activates only 6B at inference, and that the Next architecture—distinct from prior Qwen models—combines Qwen Sparse Attention with a GDN hybrid attention stack, a Gated Residual pathway, and 51B N-gram Embedding parameters that sit outside the per-token compute path. Alibaba claims training cost fell nearly 90% versus Qwen3.7-Plus, API pricing of RMB 1 per million input tokens and RMB 3 per million output tokens (as low as one-third DeepSeek-V4-Flash busy-hour pricing), and post-training scores ahead of Claude Opus 4.6 on agentic and multimodal benches including SWE-bench Pro, CoWorkBench, Toolathlon Verified, JobBench, AndroidWorld, MathVision, and ERQA. Weights are on Hugging Face and ModelScope; the model ships first in Qwen Office standard mode and via the Qwen AI platform API. The firm frames the Next stack as the prototype for Qwen4, and notes the Qwen3.8 family now spans Max (2.4T), 27B, and Flash, with Qwen downloads past 3 billion and more than 300,000 derivatives.

This is another hard move in the Chinese open-weight efficiency race: frontier-class claims at a small active-parameter budget, with list prices that undercut both Western closed models and peer CN Flash SKUs. It updates the foundation-model player map around Alibaba’s Qwen line and reinforces the open-weight compression thesis—capability per activated FLOP and per API dollar—against closed-frontier pricing. Relative to Alibaba’s recent Qwen3.8-Max monetization talk, Flash doubles down on open release as the proving ground before Qwen4.

Builders should treat the Hugging Face/ModelScope drop as an immediate eval candidate for agentic coding, long-context office workflows, and multimodal tool use where activation cost matters; investors should watch whether the claimed 90% train-cost cut and RMB-level token prices hold under third-party load and whether Next becomes the lasting Qwen4 spine or a one-off Flash SKU.

#Alibaba #Qwen #OpenWeights #FoundationModels #MoE #AI

#Alibaba#Qwen3.8-Flash#open weights#mixture of experts#Qwen4#foundation models
Read Original

How This Connects

Based on Foundation Models · Player Map

  1. 5h agoAlibaba open-weights Qwen3.8-Flash, a 125B multimodal MoE previewing Qwen4. · THIS ARTICLE
  2. 5h agoOpenAI Reports Detail a Rogue Model Collective’s Cybersecurity BreachOpenAI
  3. 21h agoMistral partners with Saudi Arabia's Humain on Arabic frontier models and regional AI infrastructure.Mistral
  4. 1w agoAlibaba's Qwen team has open-sourced Qwen3.8-27B, a 27-billion-parameter multimodal model designed f...Qwen
  5. 1w agoAlibaba has released Qwen 3.8 27B, an Apache 2.0-licensed open-weight dense model with 27 billion pa...Alibaba Qwen 3.8 27B launch
  6. 1mo agoMoonshot AI launches Kimi K3, a 2.8 trillion-parameter open-weight model, claiming performance near US frontier labsMoonshot AI

Related News

Discover AI Startups

Explore 5,000+ AI companies with VC-grade analysis, funding data, and investment insights.

Explore Dashboard