Skip to content
Friendly disclaimer: flozi00 TechHub is a solo side-project next to a full-time job — personal learning notes, no official statements. Verify critical steps yourself.

Inside Qwen3.8-Flash-Next: Gated DeltaNet, Sparse Attention, and N-gram Memory

A source-checked architecture guide to Qwen3.8-Flash-Next's released weights: the 3:1 GDN/QSA stack, gated residual, MoE routing, offloaded n-gram embeddings, and the precise overlap with DeepSeek V4.1 Flash's Engram.

5 min readflozi00
aillmarchitectureqwendeepseekengramsparse-attentionmoeinference

Qwen3.8-Flash-Next is an open-weight, multimodal model released on 26 August 2026. Its architecture is available in the Qwen team's technical report, model card, and checkpoint configuration. It combines recurrent token mixing, selective attention, routed experts, four residual branches, and a large n-gram lookup table. The released checkpoint is the subject here; Qwen describes the separate Qwen3.8-Flash API model as based on Flash-Next with additional production features, so the two names should not be treated as identical checkpoints.12

The key distinction is between three forms of memory. Gated DeltaNet (GDN) keeps a fixed-size recurrent state; Qwen Sparse Attention (QSA) retrieves selected tokens from the current context; the n-gram table holds learned parameters addressed by short token sequences. The n-gram table is neither the request's KV cache nor a router selecting MoE experts.3 The VRAM guide covers the separate capacity budget for weights and inference state.

The released stack

ComponentPublished configurationRole
Main language model125B parameters; about 6B activated per tokenSparse MoE backbone, excluding the n-gram table and MTP module.2
Token-mixing layers48 layers: 12 repetitions of three GDN layers followed by one QSA layerGDN updates compact state; QSA retains content-based access to older tokens.23
MoE512 routed experts, ten selected per token, plus one shared expertAdds conditional compute after token mixing; not the n-gram lookup.2
ResidualFour branches; gated read and per-branch scalar writeCarries different information across layers without reducing everything to one residual vector.3
N-gram memory51B additional parameters; bigram/trigram lookup at layer 2Adds learned capacity via sparse table reads.2
Draft moduleOne MTP layer, approximately 4B additional parametersSupports multi-token prediction; not included in the 125B backbone count.2

The model card gives 125B + 51B + 4B for backbone, n-gram table and MTP respectively. Its 6B active figure describes the sparse main-model computation; it does not mean the remaining stored parameters disappear from the memory budget. At two bytes per unquantized BF16 parameter, the 51B table alone represents roughly 102 GB decimal of raw weights, before allocator, sharding and caching overhead. This is arithmetic from the published parameter count and dtype, not a measured serving footprint.24

How a token moves through the model

1. Hybrid token mixing. A 48-layer stack alternates 36 GDN and 12 QSA layers. GDN writes history into a recurrent state, avoiding a per-token KV record in those layers, but a finite state cannot provide exact random access to every earlier token. Every fourth layer therefore has attention over the sequence. During continued pretraining, Qwen replaced the dense attention in those slots with QSA; the released configuration still labels those slots full_attention, while the report states that their implementation is QSA.34

2. QSA's two stages. The indexer first projects each token to lightweight keys, averages groups of four keys into micro-block keys, then scores those blocks with four query heads and one shared key head. RoPE is applied after key pooling so vectors from different positions are not averaged with different rotary phases. For each query, the reported budget is 512 selected complete blocks = 2,048 tokens, plus any incomplete tail block. The core attention then computes over those selected tokens with 24 query heads and two KV heads of dimension 256.324 Compressing the indexer by four reduces its prefill work, but does not make the complete system constant-cost with context length. Qwen's reported 7.6× prefill and 4.9× decode speedups at 1M context compare QSA with dense attention at the kernel level; they are not end-to-end model latency measurements.3

3. Gated residual. At each sublayer, the four residual branches are normalized separately. A data-dependent, elementwise gate chooses how much of each branch to read into the sublayer. The result is written back with one learned scalar per branch. This is a separate mechanism from MoE routing: the four branches control information flow between layers, while the expert router selects feed-forward weights for a token.3

4. N-gram lookup. At layer 2, short token sequences ending at the current position—bigrams and trigrams in the released model—are hashed into learned embedding rows. Qwen's checkpoint specifies a 20-million-entry n-gram vocabulary base, eight hash heads per n-gram, a 2,560-dimensional embedding, and ple_layer_ids: [2].24 The retrieved vector is contextually gated into the residual stream. Addresses depend on token IDs, so a host-resident table can prefetch rows while the first transformer layer runs.3 Sparse access lowers per-token arithmetic, but it shifts cost to table storage, host memory, transfer bandwidth and latency hiding. The report's placement ablation found no consistent gain from splitting a fixed table budget across two layers, which is why the released design uses one.3

Is this the same as DeepSeek's Engram?

The principle is shared: learned n-gram-addressed memory augments transformer activations through sparse lookup. DeepSeek V4.1 Flash's model card calls its module Engram; Qwen calls its feature N-gram Embedding or PLE in the checkpoint. They are not the same model architecture or table layout.52

Qwen3.8-Flash-NextDeepSeek V4.1 Flash
Learned lookup capacity51B n-gram parameters196B Engram parameters25
PlacementOne layer: ple_layer_ids: [2]Two modules: engram_layer_ids: [1, 14]46
Main token path3 GDN + 1 QSA per four layers; gated residual20-layer causal encoder + 20-layer decoder; CSA2 sparse attention and Single-Pass mHC35
Sparse compute512 routed experts, ten selected + one shared384 routed experts, six selected + one shared25

DeepSeek's existing architecture article explains its encoder-decoder KV compression in depth. The table here isolates the lookup-memory overlap; a higher table parameter count alone establishes neither better quality nor faster inference. The two model teams use different backbones, training data, quantization and serving paths, so their headline counts and benchmark scores do not form a controlled ablation.

What the published evidence does and does not show

Qwen released weights and gives serving instructions for Transformers, vLLM and SGLang.12 The report's ablations support the architectural choices under their stated training and evaluation conditions. For example, increasing n-gram vocabulary size reduces loss in one experiment, while downstream scores saturate or fluctuate; more table rows should not be translated into a promised application-level improvement.3 QSA's kernel numbers likewise require a workload-specific test with the chosen runtime, context distribution, quantization and offload policy. The native checkpoint context is 262,144 tokens; the model card describes extension to one million tokens with modified RoPE/YaRN settings, so that larger number is not the default checkpoint configuration.24

Sources

Footnotes

  1. Qwen team's release repository, 26 August 2026. ↩ ↩2

  2. Qwen3.8-Flash-Next official model card. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14

  3. Qwen Team, On the Design of Qwen3.8-Next Architecture, arXiv:2608.30320, 31 August 2026. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11

  4. Qwen3.8-Flash-Next released config.json. ↩ ↩2 ↩3 ↩4 ↩5 ↩6

  5. DeepSeek V4.1 Flash official model card. ↩ ↩2 ↩3 ↩4

  6. DeepSeek V4.1 Flash released config.json. ↩