DeepSeek-V4.1-Flash — KV-Cache-Compressed Inference

DeepSeek-V4.1-Flash — KV-Cache-Compressed Inference An architecture diagram generated by Archify. Images · input · Architecture component Images input Text · input · Architecture component Text input ViT + Projector · 2D-RoPE · unshuffle ×9 · Architecture component ViT + Projector 2D-RoPE · unshuffle ×9 Encoder layers 1–2 · SWA only · window 128 · Causal encoder — 20 layers Encoder layers 1–2 SWA only · window 128 Encoder CSA2 · 18 layers · m=2 · Full + 5 Reuse · Causal encoder — 20 layers Encoder CSA2 18 layers · m=2 · Full + 5 Reuse Decoder CSA2 · 20 L · m=1 · HSI · Top-512 · Architecture component Decoder CSA2 20 L · m=1 · HSI · Top-512 Tokens · autoregressive · Architecture component Tokens autoregressive SWA pool · host DRAM · minutes TTL · replay 128 tok · Tiered KV storage SWA pool · host DRAM minutes TTL · replay 128 tok Global KV · HBM · 890 B/token · MXFP4 · Tiered KV storage Global KV · HBM 890 B/token · MXFP4 Persistent · SSD · global KV ≥72 h · Tiered KV storage Persistent · SSD global KV ≥72 h images visual emb. hidden states CED projection main KV m=2 main KV m=1 SWA replay · 128 tok prefix reload Causal encoder — 20 layers Tiered KV storage Legend Backend Database External

Why decode stays cheap

  • • Global KV: 890 bytes per token in HBM — about 1/4 of DeepSeek-V4-Flash at equal context
  • • Single-token Decode FLOPs rise only ~1/4 while context grows 4K to 1M (256×)
  • • Reuse-Mode layers run 15 kernels in prefill and 11 in decode

The four compression axes

  • • CED: prefill activates 8B params, decode 16B — decoder global KV is projected from the final encoder state
  • • CSA2: main KV and Top-512 indices shared across layer groups in Full / Reindex / Reuse modes
  • • Main KV stored in MXFP4 (E2M1 + E4M3 scale per 16 channels); SWA KV stays FP8
  • • SWA Bounded Replay: reconstruct local state from the last 128 tokens, not L × 128 = 5,120

Supporting cast

  • • DSpark drafter: 3 blocks, 5 draft positions per pass, confidence-scheduled verification
  • • Engram: 196B sparse conditional memory at layers 1 and 14, FP8 tables, RDMA prefetch
  • • Single-Pass mHC + Mega-mHC kernel halves activation memory traffic of the multi-kernel variant

What to watch

  • • One Full-mode selection error poisons every Reuse layer that shares it
  • • Bounded-replay states depend on the cache-hit position — not bit-identical across resumption points
  • • FP4's 1-bit mantissa plus needle retrieval at 512K–1M context is the open stress case