Qwen-Image 2.1 makes part of a diffusion transformer's sequence immutable across denoising steps. Text and reference-image tokens form a block-causal prefix. The noisy target image can read that prefix, but the prefix cannot read the target. Prefix tokens are additionally modulated at timestep zero instead of the current noise level. Their hidden states, keys, and values therefore remain valid from one denoising step to the next.
That is the mechanism behind Qwen-Image 2.1's prefix KV cache. It is not the same lifecycle as an autoregressive LLM cache: the reused history is fixed conditioning data, while all target-image tokens are recomputed together at every diffusion step. This article covers the released base checkpoint from 20 September 2026 and the self-hostable Turbo checkpoint from 9 October 2026. Sources were checked on 10 October against immutable model and Diffusers revisions.123
The released pipeline
Qwen-Image 2.1 is a unified text-to-image and image-editing model. The same pipeline accepts a prompt, up to ten reference images, and a noisy target canvas. Its 7B visual generator is a 32-block single-stream diffusion transformer. Official integrations include Diffusers and ComfyUI, with serving paths documented for vLLM-Omni, SGLang, and LightX2V.12
The complete pipeline has three relevant stages:
- A Qwen3-VL encoder maps the prompt and reference images into 4,096-wide contextual representations.
- A VAE maps reference and target images to a 64-channel latent grid at 16Γ spatial downsampling.
- The 32-block visual transformer predicts the flow update for every target latent at the current denoising timestep.
The checkpoint uses the Qwen Research License rather than Apache 2.0. Self-hosting is technically supported, but production use still requires reviewing that license and the application's obligations.4
One stream for text, reference images, and the target
Let be the number of ordinary text positions after prompt encoding, the latent-token count of reference image , and the target-image latent count. With width , height , and VAE scale factor 16:
The VLM initially represents each reference-image slot at coarser granularity. The Diffusers implementation expands every such slot four-fold because it corresponds to a 2Γ2 group of VAE latent positions, then replaces those expanded positions with the actual 64-channel reference latents. Target latents are appended at the end. Linear input projections map text width 4,096 and latent width 64 into the transformer's shared width of 4,096.5
Each transformer block uses 32 attention heads of width 128. For a joint sequence [B, S, 4096], the unfused Q, K, and V tensors each have shape [B, S, 32, 128]. This is ordinary multi-head attention, not GQA: every query head has its own key and value head. Per-head RMS normalization is applied to Q and K before RoPE. The attention result returns to width 4,096; a SwiGLU feed-forward path expands to 12,288 and projects back.65
Block-causal attention is the cache-enabling boundary
Every position receives an image-block identifier: -1 for text, a unique non-negative identifier for each reference image, and another identifier for the target. For query position and key position , ignoring padded keys, attention is allowed exactly when:
Text is strictly causal. Every image block is internally bidirectional, so all latent positions of one image can interact in the same layer. Across blocks, sequence order remains causal. Consequently:
- a reference image can use earlier text and every position inside itself, but not later text, later images, or the noisy target;
- later prompt text can read an earlier reference image;
- the final target block can read the complete prefix and every position inside the target.
Adjacent reference images retain different identifiers even when no text separates them. Treating them as one image block would incorrectly allow the earlier image to read the later one. This mask is more than an attention optimization: it establishes the dependency boundary that makes prefix reuse mathematically valid.5
Timestep-zero conditioning makes the prefix invariant
Diffusion normally injects the sampled timestep into every token. If reference tokens were modulated by the current timestep, their layer states and K/V projections would change at every denoising step even though their pixels did not. Qwen-Image 2.1 sets causal_condition=true and creates an additional timestep embedding for .65
At every block, text and reference-image positions select scales and residual gates derived from ; target positions select those derived from the current . A single shared modulation projection produces these parameters for all 32 blocks. Combined with the attention mask, prefix states cannot depend on either or target latents:
On the first denoising step, kv_cache_mode="extract" runs the complete joint sequence and stores post-RoPE prefix K/V for every layer. Later steps use "cached": only target hidden states enter the transformer blocks, while cached prefix K/V are concatenated with freshly computed target K/V. The implementation clones the prefix slices. A mere contiguous view at batch size one could otherwise keep the complete prefill K/V allocation alive for every step.57
The cache removes repeated transformer processing of the prefix. It does not remove the target's attention over that prefix. At every step and layer, target queries still score keys, where . The distinction mirrors the general KV-cache guide: cached projection work and attention over cached history are different costs.
The prefix cache is a compute-for-memory trade
The released transformer caches two tensors for 32 layers, 32 heads, and head width 128. For batch size , prefix length , and bytes per cached scalar, the logical payload is:
At BF16 and batch one, that is 524,288 bytes = 0.5 MiB per prefix token. A 1,024Γ1,024 reference image contributes 64Γ64 = 4,096 latent tokens and therefore 2 GiB of raw prefix K/V. A 2,048Γ2,048 reference contributes 16,384 tokens and 8 GiB. Prompt text adds 0.5 MiB per final joint text position. These are tensor-payload calculations, not measured peak VRAM.68
The actual allocation also includes transformer and Qwen3-VL weights, VAE state, target latents, activations, attention workspaces, allocator rounding, and backend-specific buffers. Tensor parallelism can shard projections and cache differently. If true classifier-free guidance is enabled, Diffusers performs a positive and a negative transformer pass and creates a separate prefix cache for each; both compute and cache payload approximately double. The production default is guidance scale 1, where the negative path is absent.79
For denoising steps, uncached execution processes prefix tokens through all blocks times. Cached execution does so once, then processes only target tokens for the remaining steps. The target-to-prefix attention term remains on every step. Therefore speedup depends on , resolution, attention backend, memory bandwidth, and the cost of other pipeline stages. A cache can be valuable for editing with large reference images while simultaneously becoming a major VRAM consumer.
Three-axis RoPE preserves image geometry
The 128 dimensions of each attention head are divided across frame, height, and width axes as (16, 56, 56), with RoPE base 10,000. Text positions advance one shared coordinate on all three axes. Within an image block, the frame coordinate is fixed at the position reached by preceding text, while height and width coordinates form a grid centered around zero.65
This makes an image's spatial coordinates independent of where its block appears in the prompt. Moving a reference block later changes its frame coordinate but does not translate its internal height/width grid. RoPE changes Q and K, not V, and does not change the 128-value cache width.
Two exact attention paths, with different operating requirements
The default Diffusers processor implements the block-causal prefill as multiple exact attention calls: one for every text or image prefix segment, followed by one for the target. A text segment receives a causal triangle over itself; an image segment attends to all earlier keys and its complete own block. This works without compilation but launches multiple attention operations per layer.5
The optional FlexAttention processor represents the same rule with a BlockMask quantized to 128-token blocks and executes prefill in one call. It requires PyTorch 2.5 or newer and should be compiled. The pinned implementation warns that uncompiled FlexAttention materializes a dense FP32 score matrix and can run out of memory at high resolution. Cached later steps no longer need the block-causal mask: target rows have full access to [cached prefix, target], so the configured attention backend handles them directly.59
Backend selection is therefore operational, not a model-quality choice. Compare prefill latency, denoising-step latency, peak VRAM, compilation cost, supported resolution, and parallelism on the exact runtime revision. The same checkpoint can take materially different kernel paths.
What Turbo changesβand what it does not
Qwen-Image 2.1 Turbo has the same visual-transformer architecture and tensor shapes as the base model. Both configurations specify 32 layers, 32 heads, 128 dimensions per head, width 4,096, 64 latent channels, and causal_condition=true. Turbo is an accelerated checkpoint trained for an eight-step trajectory, not a smaller transformer.1011
The base pipeline defaults to 40 denoising steps and dynamically shifts its FlowMatch Euler schedule with sequence length. Turbo stores this eight-value sigma schedule in model_index.json and disables dynamic shifting:
1.000000, 0.978453, 0.954180, 0.926626,
0.895080, 0.845148, 0.704534, 0.414568Diffusers loads the saved schedule automatically. Setting only num_inference_steps does not replace it. Eight versus 40 means five times fewer denoiser passes by count, but not a guaranteed 5Γ wall-clock speedup: prompt/reference encoding, VAE encode/decode, first-step cache prefill, memory movement, and runtime overhead do not shrink by that factor.37
Prefix caching can also change output bits in reduced precision. Diffusers documents that cached and uncached paths tile the same computation differently; a one-ULP early difference can amplify through 32 blocks and multiple steps. Both agree with the FP32 reference within the same tolerance, but a reproducible deployment must pin the checkpoint, schedule, precision, backend, compilation state, and use_kv_cache flag.
Production boundaries
The architecture is useful because it separates immutable conditioning from the time-dependent target without giving up bidirectional interaction inside each image. The benefit is conditional: prefix reuse saves repeated prefix computation, while high-resolution references create a large persistent K/V allocation. Turbo reduces step count, while leaving per-step transformer width and target attention unchanged.
For capacity planning, calculate target and reference latent counts from actual post-resize dimensions, include one positive and optionally one negative cache, then measure peak VRAM through prefill and later steps. Keep cache precision separate from weight quantization. Benchmark base and Turbo with identical prompts, seeds, resolutions, references, backends, and output criteria; eight-step generation is a different checkpoint and trajectory, not the 40-step model with a loop counter changed.
Primary sources and revision boundary
Model files and implementation links below are pinned. The Qwen release repository is pinned at its 9 October announcement revision. Diffusers is pinned at the revision used for the implementation analysis. Recheck cache layout and backend behavior when upgrading.
Footnotes
-
Qwen, Qwen-Image 2.1 release repository at
1993c3a: release dates, integrations, and Turbo usage. β© β©2 -
Qwen, Qwen-Image 2.1 Turbo checkpoint at
d65dbc9: eight-step schedule and usage. β© β©2 -
Qwen, Qwen-Image 2.1 Turbo license. β©
-
Hugging Face,
QwenImage21Transformer2DModelat Diffusers1d5d056: sequence construction, mask, RoPE, blocks, processors, and cache. β© β©2 β©3 β©4 β©5 β©6 β©7 β©8 -
Hugging Face,
QwenImage21Pipelineat the same Diffusers revision: schedule, CFG, and denoising-loop cache lifecycle. β© β©2 β©3 -
Qwen, base VAE configuration. β©
-
Hugging Face, Qwen-Image 2.1 Diffusers documentation at the same revision: default steps and compiled FlexAttention requirements. β© β©2
-
Qwen, base pipeline index and scheduler configuration. β©
-
Qwen, Turbo transformer configuration and pipeline index. β©