StorySoftware
Alibaba releases open-weight Qwen3.5 and proprietary Qwen3.5-Plus
Alibaba released open-weight Qwen3.5 led by a 397B-parameter mixture-of-experts model with 17B active per token under Apache 2.0, alongside proprietary Qwen3.5-Plus. The models handle text, images and video in 201 languages with 262K-token context. They combine a newer, cheaper attention design with sparse experts and native support for extended context up to 1M tokens.
- 397B-parameter open-weight model with 17B active per token under Apache 2.0; proprietary Qwen3.5-Plus also released
- Multimodal processing of text, images and video in 201 languages with 262K native context, extensible to 1M tokens
- Architecture combines gated DeltaNet and attention layers with sparse mixture-of-experts (512 experts, 11 activated)
- Built-in tool-calling and thinking mode for agentic reasoning; strong performance on mathematics (HMMT Feb: 94.8%), coding (SWE-bench: 76.4%)