Kimi K3 is a released, open-weight multimodal model, not just an architecture proposal. Moonshot published its weights, model card, checkpoint configuration, and technical report in July–August 2026. The model card documents API access and deployment recipes. This article describes the released moonshotai/Kimi-K3 checkpoint; the team's performance claims are not independent serving measurements.123
Three axes matter: Kimi Delta Attention (KDA) compresses token history into a recurrent state, periodic Gated Multi-head Latent Attention (MLA) retains global token access, and Attention Residuals (AttnRes) select information from earlier layers. Stable LatentMoE sparsifies the feed-forward path across experts. None is a learned n-gram table: the Qwen3.8-Flash-Next article explains that different kind of parameter memory.3
Checkpoint anatomy
| Part | Published value | Consequence |
|---|---|---|
| Backbone | 93 attention layers, hidden width 7,168; 2.8T total and 104B activated parameters | Active compute and stored weights have different budgets.1 |
| Sequence mixer | 69 KDA and 24 Gated MLA layers | Layers 1–92 follow three KDA then one MLA; layer 93 is an additional MLA layer.32 |
| Routed feed-forward | Latent width 3,584; 896 routed experts, 16 chosen per token, plus two full-width shared experts | Routing communicates a half-width token vector to selected experts.13 |
| Depth mixer | Block AttnRes with 12 layers per block | Eight partial/full layer blocks, plus the embedding as a depth source.32 |
| Vision | MoonViT-V2, 27 layers and about 401M parameters | Visual features are projected into the shared language backbone.31 |
| Context | max_position_embeddings: 1048576 | A supported context ceiling, not a guarantee of cheap one-million-token requests.2 |
The configuration marks one initial dense feed-forward layer (first_k_dense_replace: 1). The 93 attention slots still divide into 69 KDA and 24 MLA. Importantly, KDA's recurrent state, MLA's per-token KV, and AttnRes's earlier-layer representations are three distinct runtime states.23
Sequence mixing: fixed-state KDA alongside global MLA
For one KDA head, the report defines query and key vectors of dimension , value vectors of dimension , and a recurrent matrix . At each token, a channel-wise retention vector decays the old state; a scalar write gate applies a delta-rule correction before the query reads . This is why a KDA layer need not store a KV vector for every previous token. The released configuration gives 96 heads of dimension 128 and a short convolution of width 4 for the KDA projections. The state is fixed-size with respect to sequence length, but it does not provide exact, unrestricted retrieval of every old token.32
For prefill, KDA evaluates tokens in parallel inside a chunk while passing the recurrent state between chunks. Its causal intra-chunk matrix includes the diagonal because each position reads the state after its own update. Numerical range matters: the published log-decay is bounded by , so a one-step retention factor exceeds . For a 16-token tile, the cumulative log-decay remains above ; the report says this permits the diagonal as well as off-diagonal tiles to use Tensor Core matrix multiplication instead of an explicit position-pair path. This is a kernel-enabling bound, not a bound on information loss across a million tokens.32
Every fourth layer through layer 92, and again at layer 93, uses Gated MLA. MLA caches a compressed token representation rather than full head-specific keys and values; the checkpoint sets its KV latent rank to 512. Unlike Kimi K2, the K3 report says the MLA path uses NoPE, with no explicit positional encoding on its queries or keys. KDA supplies position-sensitive mixing between those global layers. Both KDA and MLA have input-dependent output gates, but their gates act on their own attention outputs; they are separate from the MoE router.32
The combination therefore does not have constant memory per request: 69 KDA states stay fixed-size while 24 MLA caches grow with retained context. For prefix reuse, the recurrent KDA state and MLA KV must be restored at the same token boundary. The report describes a shared paged cache and sparse KDA checkpoints; the latter limit which prefixes can be reused without recomputation. The VRAM guide covers weight and KV budgets more generally.3
Depth mixing: Block Attention Residuals
Ordinary residual addition carries previous layers forward in one running vector. AttnRes instead scores earlier representations with a learned, layer-specific query, applies RMSNorm to each candidate before scoring, and softmax-normalizes the weights across depth, per token. Its full form would retain every layer output. K3 uses Block AttnRes: outputs inside each 12-layer block are summed; later layers attend to completed block summaries, the embedding, and the current block's partial sum. The report describes eight layer blocks and nine sources when the embedding is counted.32
This is depth attention, not attention across earlier token positions. The block form reduces the number of stored depth summaries from order to order , where is layer count and block count; it still requires the current block's partial sum and its normal inference activations. It does not remove MLA's sequence KV cache.3
Width mixing: 16-of-896 Stable LatentMoE
For a hidden vector , the routed branch first projects to . A sigmoid router selects 16 of 896 expert networks; their outputs are weighted and summed at latent width. RMSNorm then precedes an up-projection back to width 7,168. Two shared experts process the full-width input directly, and their outputs are added to the routed branch. The reported per-expert intermediate width is 3,072. The expert choice is token-dependent; 104B "activated parameters" does not mean the other weights can be absent from a serving deployment.312
The feed-forward activation is SiTU-GLU, which soft-caps both multiplicative branches with scaled tanh functions. The published caps are 4 and 25, yielding an elementwise product bound of 100 before downstream projections. Quantile Balancing adjusts expert-selection biases during training based on batch load; the report says these biases are frozen at inference. The router uses the biases to choose experts but uses the original sigmoid scores to weight their outputs. These are concrete stabilization and routing rules, not evidence that any particular deployment will have balanced per-request expert traffic.32
Weights, multimodality, and operational limits
The model card calls the model native MXFP4-weight / MXFP8-activation quantization-aware trained. The released config.json is more specific: its compressed-tensors group targets certain Linear weights in MXFP4 with group size 32, while the ignore rules exclude attention, shared experts, the language-model head, the vision tower, the multimodal projector, and some other projections. Thus "2.8T parameters × four bits" is not a sound checkpoint-size or device-memory estimate. Quantization metadata also does not specify a universal runtime activation dtype for every kernel.12
MoonViT-V2 encodes visual input before an MLP projector feeds the language backbone. The report discusses images and video; the model-card summary lists text and image, while its introduction also mentions video. Support for a particular video input format depends on the published processor and serving interface, so a generic text/image example should not be taken as proof of every video path.31
Moonshot provides model weights and recommends vLLM, SGLang, and TokenSpeed for deployment; the model card also documents the kimi-k3 API. A real installation still has to budget all expert weights, mixed-precision tensors, 24 length-dependent MLA caches, KDA state/checkpoints, and MoE communication. The team's benchmark and scaling-efficiency figures are results under its own test conditions, not throughput guarantees. A useful comparison is the DeepSeek V4.1 Flash architecture article, whose cache design makes different storage and replay trade-offs.13
Sources
Footnotes
-
Moonshot AI, Kimi K3 official model card and deployment instructions. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8
-
Moonshot AI, released Kimi K3
config.json. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 -
Kimi Team, Kimi K3: Open Frontier Intelligence, arXiv:2607.24653v2, 7 August 2026. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14 ↩15 ↩16 ↩17