
Alibaba open-weights Qwen3.8-Flash, a 125B multimodal MoE previewing Qwen4.
The AMW Read
Known Qwen player ships a new Next MoE architecture with open weights and aggressive cost claims, advancing CN open-weight efficiency without resolving a named debate.
Alibaba open-weights Qwen3.8-Flash, a 125B multimodal MoE previewing Qwen4.
On August 26, Alibaba released and open-sourced Qwen3.8-Flash, a multimodal mixture-of-experts model under its Qwen line. The company says the Transformer uses 125B parameters but activates only 6B at inference, and that the Next architecture—distinct from prior Qwen models—combines Qwen Sparse Attention with a GDN hybrid attention stack, a Gated Residual pathway, and 51B N-gram Embedding parameters that sit outside the per-token compute path. Alibaba claims training cost fell nearly 90% versus Qwen3.7-Plus, API pricing of RMB 1 per million input tokens and RMB 3 per million output tokens (as low as one-third DeepSeek-V4-Flash busy-hour pricing), and post-training scores ahead of Claude Opus 4.6 on agentic and multimodal benches including SWE-bench Pro, CoWorkBench, Toolathlon Verified, JobBench, AndroidWorld, MathVision, and ERQA. Weights are on Hugging Face and ModelScope; the model ships first in Qwen Office standard mode and via the Qwen AI platform API. The firm frames the Next stack as the prototype for Qwen4, and notes the Qwen3.8 family now spans Max (2.4T), 27B, and Flash, with Qwen downloads past 3 billion and more than 300,000 derivatives.
This is another hard move in the Chinese open-weight efficiency race: frontier-class claims at a small active-parameter budget, with list prices that undercut both Western closed models and peer CN Flash SKUs. It updates the foundation-model player map around Alibaba’s Qwen line and reinforces the open-weight compression thesis—capability per activated FLOP and per API dollar—against closed-frontier pricing. Relative to Alibaba’s recent Qwen3.8-Max monetization talk, Flash doubles down on open release as the proving ground before Qwen4.
Builders should treat the Hugging Face/ModelScope drop as an immediate eval candidate for agentic coding, long-context office workflows, and multimodal tool use where activation cost matters; investors should watch whether the claimed 90% train-cost cut and RMB-level token prices hold under third-party load and whether Next becomes the lasting Qwen4 spine or a one-off Flash SKU.


