HomeFull speed ahead. Keep track of advancements in every frontier
Qwen modelsAlibaba Qwen—Qwen modelsAlibaba Qwen—Qwen modelsAlibaba Qwen—
Qwen models
Courtesy of Alibaba Qwen · Press kit, editorial use (opens qwen.ai)

Qwen models

by Alibaba QwenCN

ScalingStage 5 of 5

Qwen3.8 (August 2026) is current; Alibaba said in September that Qwen4 is in training.

Updated 18d agoChecked 25 Sep0 updates this week

Milestones

Next · Qwen4 release
  1. Tongyi Qianwen unveiledApr 2023Complete.
  2. First open-weight Qwen-7BAug 2023Complete.
  3. Qwen2.5 familySep 2024Complete.
  4. Qwen3 hybrid-reasoning family29 Apr 2025Complete.
  5. Qwen3-Max announced at Apsara conference24 Sep 2025Complete.
  6. Qwen4 releaseNowCurrent milestone.

Most important updates

  • 18d ago
  • 3 Aug 2026
  • Feb 2026
  • 24 Sep 2025
  • 29 Apr 2025

Upcoming

  1. Q4 2026Qwen-Image-3.1, which Alibaba said at Apsara will launch before the end of 2026 (next)

Current obstacles

  • Access to advanced chipsUS export limits on NVIDIA chips push Alibaba toward its own and other Chinese chips for training.

Physics limits

  • Running out of human-written textFrontier models already train on 15 to 40 trillion tokens. Estimates put the usable stock of public human text at a few hundred trillion, so data, not chips, starts to cap plain scaling.
  • Shrinking numbers loses informationRunning on smaller hardware means storing weights in 4 to 8 bits instead of 16. Below about 4 bits accuracy drops quickly, so model size sets a floor on memory needed.
  • Returns shrink as a power lawError falls only as a small power of training compute: each fixed step of improvement needs roughly 10x more compute, energy and money than the one before.

How it works

4 parts
Rows of server racks in a data centre: Qwen's largest models are served from Alibaba Cloud halls of this kind
Rows of server racks in a data centre: Qwen's largest models are served from Alibaba Cloud halls of this kindPhoto: PiDatacenters · CC BY-SA 4.0 (opens commons.wikimedia.org)
Range

Every size, mostly open

Qwen ships dense models from under 1B up to 235B-parameter mixtures of experts, most under Apache 2.0, so it is a common base for others.

Hybrid

Thinking on demand

Qwen3 models can answer instantly or reason at length within the same model, switched by the user or by a thinking budget.

Data

36 trillion tokens

Qwen3 was pretrained on about 36 trillion tokens covering 119 languages and dialects, heavy in code and maths.

Max

Largest models via cloud

The biggest models, such as Qwen3-Max with over a trillion parameters, are served through Alibaba Cloud rather than released as weights.

Update log

7 updates

Tue 22 Sep

  • Major: PressResearch

Mon 3 Aug

  • Minor: OtherSoftware

Feb 2026

  • Minor: OtherSoftware

Wed 24 Sep 2025

  • Minor: PressSoftware

Tue 29 Apr 2025

  • Major: BlogSoftware

Tue 28 Jan 2025

  • Minor: BlogSoftware

Thu 19 Sep 2024

  • Minor: BlogSoftware

About Alibaba Qwen

The team behind Qwen models

Alibaba Qwen

Qwen open-weight models, Qwen app

Qwen is Alibaba Cloud's AI model family and one of the most used open-weight lines, with hundreds of thousands of spin-off models online. Sizes range from under 1 billion to over 1 trillion parameters.

  • Qwen3 (April 2025) can answer fast or think step by step, and is open under a permissive licence.
  • Alibaba committed RMB 380B (about $53B) over three years to cloud and AI in February 2025.
  • Qwen3-Max, the closed flagship, has over a trillion parameters.
Founded
20233 yrs
Headquarters
China
Status
Subsidiaryof Alibaba Group
Valuation
Alibaba Group-owned
Works in
AIOpen-weight models
Coverage
1 program · 8 updateslatest 18d agochecked 25 Sep
Primer