Blog / Qwen 4
Qwen 4 announced: what we know so far
Alibaba announced Qwen 4 at its Apsara conference in Hangzhou on 22 September. The company says the new generation is still in training.
If you want to run Qwen4 on your own hardware, the 27B model is worth watching. Qwen released Qwen3.6-27B in April and Qwen3.8-27B in August as open weights under Apache 2.0, which allows commercial use without a separate agreement. Our guess is that Qwen4-27B will be released the same way.
The other three models are unlikely to be downloadable at launch. Alibaba offered Qwen3.6-Plus, Qwen3.7-Max and Qwen3.7-Plus only as hosted models on Alibaba Cloud. Qwen3.8-Max's weights came out in August as Qwen3.8-2.4T-A95B, under a custom licence. If Qwen4 follows the same pattern, the 27B model will be downloadable first, Max may follow under similar terms, and Plus will stay API-only.
Qwen gave an early look at the new architecture behind Qwen4 in August, when it released Qwen3.8-Flash-Next. It accepts text, images and video, and can handle about 262,000 tokens of context, with an option to extend that to a million.
The preview's main network has 125 billion parameters, with about 6 billion active for each token. It also has a 51-billion-parameter lookup table for short token sequences, which can be kept in system RAM instead of GPU memory. Fewer active parameters cut computation. All the weights still have to be loaded into memory, but much of that can be system RAM, which is far cheaper than GPU memory. The trade-off is fewer tokens per second.
Qwen reports three efficiency gains for the Flash-Next architecture:
- Training: about one ninth of the cost of Qwen3.7-Plus.
- Processing prompts: 8.6 times the throughput of Qwen3.7-Plus when 90% of each prompt is already in the prefix cache from earlier requests.
- Handling long inputs: at one million tokens, the Qwen Sparse Attention kernel is up to 7.6 times faster at processing input and 4.9 times faster at generating responses. The speed-up applies to the attention step only, not the whole model.
The team also says the base preview beat Qwen3.7-Plus-Base on 8 of 14 benchmarks while using less computation.
Flash-Next uses the Qwen Community License 1.0, not Apache 2.0. If Qwen4-Flash gets open weights, we'd expect the same licence.
At the same conference, Alibaba described an experiment in which Qwen3.8-Max helped improve its own training. The company says the model ran 33 automated rounds in a month, raising its Artificial Analysis score from 40 to 45.
Alibaba also outlined plans for Qwen 4.5 and Qwen 5 models with 5 to 10 trillion parameters, but gave no parameter count for Qwen4-Max.
There are no Qwen4 prices or benchmarks yet, so it's too early to say how the models compare with what you can use today. We expect the biggest gains in long prompts, coding and multi-step tasks. If the preview's efficiency gains carry over, Plus and Flash could offer better performance for the same price.