Section 24

Kimi K2.5

Trillion-parameter MoE and the Muon optimizer

Listen to this chapter

Paper: Kimi K2.5: Visual Agentic Intelligence — Moonshot AI, 2026

Kimi K2.5 (Moonshot AI, Feb 2026) is a multimodal model designed to take actions over several steps, but its pre-training contribution is a sharp, surprising finding about how to fold vision into a language model from the start.

Built on a trillion-parameter MoE base

Kimi K2.5 is built on Kimi K2, a roughly trillion-parameter Mixture-of-ExpertsMixture of ExpertsMixture of Experts (MoE) — a layer with many parallel sub-networks ("experts") where a router sends each token to only a few. The model has a huge total parameter count but activates only a fraction per token, so compute stays modest.See in glossary → transformer. As in DeepSeek-V3, routing tokens to selected experts allows a large total parameter count while limiting the work per token.

Muon without and with the qk-clip guardrail
Raw Muon lets attention logits blow up mid-run; MuonClip rescales the query and key projections before they can.
Raw Muon
logit blow-upslosstraining steps

Orthogonalized updates keep growing a few query-key logits until softmax saturates and the loss spikes, again and again.

MuonClip
smooth descentlosstraining steps

Same optimizer, same run, but logits are kept bounded, so the trillion-parameter pre-training proceeds without a single spike.

The qk-clip step
Watch: after each update, track the largest query-key logit in every attention head.
Trigger: if a head's max logit exceeds the threshold t, flag that head for rescaling.
Rescale: shrink its query and key weights so the product drops back below t; unaffected heads are untouched.
Both loss curves are stylized illustrations, not measured data. Redrawn after Section 4.1 of Kimi K2.5: Visual Agentic Intelligence (Moonshot AI, 2026).

The new idea: native multimodal pre-training

The conventional way to make a language model see is to train a strong text model first, then bolt on vision late in training by adding visual tokens. Kimi K2.5 rejects this. It does native multimodal pre-trainingnative multimodal pre-trainingTraining on a mix of text and other modalities (e.g. vision) from the very start, with a constant ratio, rather than bolting a modality onto a finished text model late in training. Kimi K2.5's approach.See in glossary →: text and vision tokens are mixed at a constant ratio throughout the entire run, with vision fused in early rather than late. A vision encodervision encoderA module (such as SigLIP) that converts an image into a sequence of embedding vectors the language model can attend to, as if they were tokens. The bridge that makes a text model multimodal.See in glossary → (MoonViT-3D, a native-resolution encoder) turns images into tokens that join the stream from the beginning.

K2.5 was pre-trained this way on roughly 15 trillion mixed visual-and-text tokens.

This is the same underlying move as Gemma 3’s vision tokens, taken to its logical end: don’t adapt a text model to see. Pre-train a model that sees and reads at once. It’s also the philosophy that the omni-modal models (chapter 27) push across every modality. Its later training for taking actions and coordinating multiple agents is discussed in the post-training explainer.

Coding models face a related data-design problem: useful examples often need to be constructed around runnable code and checks of its behavior.