Qwen3.8-Flash-Next released with 125B total parameters and 6B activation per token
On September 9, Alibaba's Tongyi Qianwen team released Qwen3.8-Flash-Next, with 125B total parameters and only 6B activation per token. Doubling the batch size on a 10.8B validation model reduced training loss by 0.0072.
On September 9, Alibaba's Tongyi Qianwen team released the Qwen3.8-Flash-Next model. The model has 125B total parameters but only 6B activation per token, far less than the activation scale of the previous-generation 397B flagship.
The team validated the training strategy on a 10.8B model: doubling the batch size reduced training loss by 0.0072, making training more stable. This result supports the team's research direction of expanding model capacity without significantly increasing activation parameters. Compared with the previous generation, Qwen3.8-Flash-Next is further optimized for coding, long-context reasoning, and agent scenarios.
On the product side, Alibaba simultaneously launched teacher and student memberships for the Qianwen app. After identity verification, the Premium plan is ¥9.9/month and the Elite plan ¥25/month, for six consecutive months, primarily targeting high-frequency work scenarios such as lesson preparation and research. Qianwen Office launched a multi-person workspace on September 7, and on September 8 officially opened support for web pages generated for up to 100 collaborators.