HomeFull speed ahead. Keep track of advancements in every frontier
Moonshot AI—Kimi modelsMoonshot AI—Kimi modelsMoonshot AI—Kimi modelsMoonshot AI—Kimi modelsMoonshot AI—Kimi modelsMoonshot AI—
Kimi models
Courtesy of Moonshot AI · Press kit, editorial use (opens kimi.ai)

Kimi models

by Moonshot AICN

ScalingStage 5 of 5

Kimi K3, released July 2026, is Moonshot's largest open-weight model.

Updated 27 Jul 2026Checked 25 Sep0 updates this week

Milestones

No announced next step
  1. Kimi chatbot with 200K-character contextOct 2023Complete.
  2. Kimi K2 open-weight 1T MoE11 Jul 2025Complete.
  3. Kimi K2 Thinking6 Nov 2025Complete.
  4. Kimi K2.5 multimodal agent modelJan 2026Complete.
  5. Kimi K3Jul 2026Complete.

Most important updates

  • 27 Jul 2026
  • 16 Jul 2026
  • Apr 2026
  • Jan 2026
  • 11 Jul 2025

Current obstacles

  • Stable training at huge scaleMoonshot built MuonClip to stop training blowing up at trillion scale. Staying stable at larger sizes is ongoing work.

Physics limits

  • All experts must sit in memoryMixture-of-experts saves arithmetic, not memory: a 1-trillion-parameter model needs about 1 TB at 8-bit precision even if only 3% runs per word, so it needs many linked chips.
  • Errors compound over long tasksIf each step succeeds 99% of the time, a 100-step task succeeds only about 37% of the time (0.99^100). Agents need near-perfect per-step reliability to work for hours.
  • Attention cost grows with length squaredIn standard attention every token is compared with every other, so doubling the input quadruples that work. Very long contexts need shortcuts that can miss details.

How it works

4 parts
GPU servers stacked in a rack: Kimi's trillion-parameter models are trained on NVIDIA GPU clusters built from machines like these
GPU servers stacked in a rack: Kimi's trillion-parameter models are trained on NVIDIA GPU clusters built from machines like thesePhoto: ChrisDag · CC BY 2.0 (opens commons.wikimedia.org)
Scale

Trillions, used sparingly

Kimi K2 has 1 trillion parameters with 32B active per word; K3 raises that to 2.8 trillion with 104B active.

Optimizer

Stable training

Moonshot's MuonClip optimizer caps runaway attention values, letting K2 train on 15.5 trillion tokens without loss spikes.

Agentic

Trained on tool use

K2 was post-trained on large volumes of synthetic tool-use tasks with RL; K2 Thinking can chain hundreds of tool calls in one job.

Context

Very long inputs

Kimi started with long context (200K characters in 2023); K3 reads about 1 million tokens at once.

Update log

7 updates

Mon 27 Jul

  • Minor: OtherSoftware

Thu 16 Jul

  • Major: BlogSoftware

Apr 2026

  • Minor: OtherSoftware

Jan 2026

  • Minor: OtherSoftware

Thu 6 Nov 2025

  • Minor: OtherSoftware

Fri 11 Jul 2025

  • Major: OtherSoftware

Mon 20 Jan 2025

  • Minor: PaperSoftware

About Moonshot AI

The team behind Kimi models

Moonshot AI

Kimi assistant, Kimi K2/K3 open models

Moonshot AI is a Beijing startup founded in 2023 by Yang Zhilin that makes the Kimi assistant. Its trillion-parameter open-weight Kimi K2 models are among the strongest open models for coding and agents.

  • Kimi K2 (July 2025) has 1 trillion parameters, 32B active, and open weights.
  • Raised about $3.5B at a $35B valuation in July 2026 after releasing Kimi K3.
  • Alibaba is a major investor.
Founded
20233 yrs
Headquarters
China
Status
Private
Valuation
$35BprivateJul 2026
Raised
$3.5B1 round
Last round
Venture round · $3.5BJul 2026$35B post
Works in
AIOpen-weight models
Coverage
1 program · 8 updateslatest 27 Jul 2026checked 25 Sep
People
Yang ZhilinFounder and CEO
Primer