Skip to content
Friendly disclaimer: flozi00 TechHub is a solo side-project next to a full-time job — personal learning notes, no official statements. Verify critical steps yourself.

Inside GLM-5.3-Flash: KDA, Sparse MLA, mHC, and 8-of-288 MoE

A source-checked guide to the released GLM-5.3-Flash checkpoint: its 3:1 KDA/sparse-MLA stack, four-stream mHC residuals, MoE routing, cache shapes, FP8 weights, and serving limits.

7 min readflozi00
aillmarchitectureglmlinear-attentionsparse-attentionmoeinference

GLM-5.3-Flash is a released, natively multimodal open-weight model from Z.ai. The team published the default FP8 checkpoint, a separate BF16 checkpoint, and deployment paths for Transformers, vLLM, SGLang and other engines on 26 August 2026. The model has approximately 320 billion parameters, activates a vendor-reported 18 billion per token, and declares a one-million-token context limit.123

It is important not to mix up the names. GLM-5.3 reuses the GLM-5.2 base model and changes post-training. GLM-5.3-Flash is a newly trained 320B base with a different glm5_next architecture. Its model card still cites the earlier GLM-5 report, which predates this Flash architecture. The low-level description below therefore relies on the released configuration and executable implementation for new 5.3-Flash details; it does not treat benchmark copy as an architecture specification.2345

The released stack

ComponentCheckpoint valueRuntime meaning
Language backbone45 layers, hidden width 4,09634 KDA layers and 11 sparse-MLA layers.4
Layer cadenceThree KDA layers followed by one sparse-MLA layer; layer 45 is an additional KDA layerMost layers keep fixed-size recurrent state; every fourth layer through layer 44 can retrieve selected earlier tokens.4
Residual pathmHC with four streams and 20 Sinkhorn iterationsEach token carries four 4,096-wide residual streams between sublayers.45
Feed-forward pathFirst three layers dense; remaining 42 layers sparse MoEEight of 288 routed experts plus one shared expert process each token in an MoE layer.45
Sparse attention budget2,048 selected token slots, index pooling factor fourThe indexer scores pools, expands selected pools to raw token indices, and retains the incomplete tail.45
Vision path24 layers, width 1,024, patch size 14, temporal patch size twoImage/video features are projected to the 4,096-wide language stream.4
Draft pathOne MTP layerRuntimes may use it for speculative decoding; the target model still verifies proposed tokens.46

These mechanisms operate on different axes. KDA and sparse MLA mix information across token positions. mHC mixes four residual streams across depth. MoE selects feed-forward weights across model width. Calling all three “routing” would hide materially different state and communication costs.

Sequence mixing: 34 fixed-state KDA layers

The linear-attention layers implement Kimi Delta Attention (KDA). For each head, the recurrent state is a matrix rather than a per-token KV list. Query, key and value pass through a causal depthwise convolution of width four. A channel-wise forget gate decays the state, a scalar β\beta controls the delta-rule write, and the query reads the updated state. The released model uses 64 heads of dimension 128 and a lower-bounded log-decay of −5-5.457

Prefill and decode follow different kernel shapes. Prefill can evaluate tokens chunkwise in parallel while carrying a state between chunks; single-token decode applies a recurrent update. In the reference Transformers implementation, one logical KDA state contains

64×128×128=1,048,57664 \times 128 \times 128 = 1{,}048{,}576

scalars per layer. Across 34 KDA layers that is 35,651,584 logical state scalars per request. The reference cache stores the recurrent state in FP32, which corresponds to about 136 MiB before convolution state, sharding and allocator overhead. This is arithmetic from the published dimensions and reference code, not a measured vLLM allocation: optimized runtimes can shard, pack or represent the state differently.45

The state size does not grow with sequence length, but “fixed-size” does not mean free or lossless. Old token information is compressed into finite matrices. The periodically interleaved sparse-attention layers provide direct retrieval for selected earlier positions. This is the same broad hybrid principle as Kimi K3, but the layer counts, hidden widths, residual design, sparse selector and expert layout differ.

Global retrieval: 11 NoPE sparse-MLA layers

Every fourth layer from zero-based indices 3 through 43 uses DeepSeek Sparse Attention (DSA) on top of Multi-head Latent Attention. The main path first compresses each token to a 512-dimensional KV latent. The checkpoint uses 64 query heads, query LoRA rank 1,536, 256 content dimensions per query/key head and 256 value dimensions. qk_rope_head_dim is zero and mla_use_nope is true: these sparse MLA layers apply no RoPE component. Causality and the intervening KDA layers still make the complete model order-sensitive.458

The separate indexer decides which earlier tokens reach core attention:

  1. It projects each query into 32 heads of dimension 128.
  2. It builds index keys of dimension 128 and compresses groups of four token keys into pools with learned gates and an intra-pool position embedding.
  3. Each query scores all visible pools, combines the 32 head scores with query-dependent weights, and selects at most 2048/4=5122048/4=512 pools.
  4. Selected pools expand back to at most 2,048 raw token indices. The current incomplete pool contributes up to three additional tail positions.

Core attention therefore works on a bounded selected set, but the indexer is not constant-time. For a prefill of length LL, it still scores queries against roughly L/4L/4 pools, so pooling reduces the indexer matrix rather than eliminating its sequence-length dependence. The index-key cache also grows with context. All entries in the released indexer_types array are full; unlike GLM-5.2's IndexShare arrangement, the 5.3-Flash checkpoint does not declare cross-layer reuse of top-k indices.45

The MLA cache itself stores 512 latent scalars per token per sparse layer and no separate RoPE vector. Across 11 layers, that is 5,632 logical KV scalars per retained token: approximately 11 KiB/token in BF16 or 5.5 KiB/token at one byte per scalar, before the indexer cache, block metadata and alignment. These figures are derived storage floors, not an end-to-end cache allocation. The KV-cache guide covers the distinction between logical tensors and runtime capacity.

Depth mixing: four-stream mHC

Manifold-Constrained Hyper-Connections replace the single residual vector with four streams. The embedding [B,S,4096][B,S,4096] is initially duplicated to [B,S,4,4096][B,S,4,4096]. Before each attention and feed-forward sublayer, the implementation flattens and normalizes the four streams, then predicts three token-dependent coefficient sets:

  • pre has four coefficients and collapses the streams to one sublayer input;
  • post has four coefficients and expands the sublayer output back to four streams;
  • comb is a 4×44\times4 matrix that mixes the residual streams.

The comb matrix is iteratively row- and column-normalized with 20 Sinkhorn-Knopp steps, constraining it toward a doubly stochastic matrix. After the sublayer, the expanded result is added to the mixed residual streams. At the end of all 45 layers, the four streams are averaged and RMS-normalized. mHC therefore increases activation traffic and coefficient work even though each attention or MLP still receives one 4,096-wide vector. Its purpose is richer and better-conditioned information flow across depth; the checkpoint alone does not prove an application-level quality gain.945

Width mixing: eight of 288 routed experts

After the first three dense layers, each feed-forward block has 288 routed experts, eight selected per token, and one shared expert. Each routed expert has intermediate width 2,048. The reference router computes its logits in FP32, applies sigmoid scores, adds a learned correction bias only for expert selection, gathers the original sigmoid scores for the chosen experts, normalizes them, and multiplies them by the configured scaling factor 2.5. The shared expert always runs and is added to the routed result.45

The expert MLP uses SwiGLU with clamping: the gate is capped above at 10 and the up branch is clamped to [−10,10][-10,10]. Sparse activation lowers arithmetic per token, but all 288 expert weight sets still have to reside somewhere accessible. Expert parallelism also turns token routing into all-to-all communication unless the runtime and placement keep selected experts local. “18B active” is therefore a compute description, not a device-memory estimate.

FP8 checkpoint and production limits

The default zai-org/GLM-5.3-Flash repository is the native FP8 checkpoint; GLM-5.3-Flash-BF16 is separate. The FP8 configuration declares dynamic E4M3 activation quantization but excludes many modules, including the attention paths, hyper-connections, embeddings, language head and selected norms/projections. The repository consequently contains BF16, FP8 and FP32 tensors. Multiplying 320B by one byte is not a complete residency calculation.34

The official vLLM recipe estimates about 306 GiB for the FP8 weights alone, before runtime state and caches, and requires a sparse-NoPE MLA backend. Its examples use TP4, optional FP8 KV on Blackwell, MTP speculative decoding, and matching KDA/KV layouts when prefill and decode are disaggregated. Those are implementation-specific deployment recipes, not universal hardware minima. A production capacity plan must include model shards, MoE communication, four-stream activations, KDA recurrent and convolution states, sparse-MLA KV, indexer state, MTP workspace and the actual request-length distribution.6

Sources

Footnotes

  1. Z.ai, GLM-5.3-Flash: Frontier Intelligence, Flash Cost, 26 August 2026. ↩

  2. Z.ai, official GLM-5 repository and release table. ↩ ↩2

  3. Z.ai, official GLM-5.3-Flash model card and released weights. ↩ ↩2 ↩3

  4. Z.ai, released GLM-5.3-Flash config.json. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14 ↩15

  5. Hugging Face Transformers, modeling_glm5_next.py. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10

  6. vLLM, official GLM-5.3-Flash deployment recipe, updated 18 September 2026. ↩ ↩2

  7. Kimi Team, Kimi Linear: An Expressive, Efficient Attention Architecture, arXiv:2510.26692. ↩

  8. DeepSeek-AI, DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models, arXiv:2512.02556. ↩

  9. Xie et al., mHC: Manifold-Constrained Hyper-Connections, arXiv:2512.24880. ↩