Skip to content
Friendly disclaimer: flozi00 TechHub is a solo side-project next to a full-time job β€” personal learning notes, no official statements. Verify critical steps yourself.

Inside Cloudflare Clef: Prefill-Only Qwen and a Joint Schema Head

A source-checked guide to the released Clef and Clef-flash decision models: hybrid Qwen backbones, option evidence routing, joint field attention, lexical scoring, tensor shapes, and production limits.

6 min readflozi00
aillmarchitecturecloudflareqwenattentionclassificationinference

Clef and Clef-flash are released multimodal decision models from Cloudflare, not research-only architecture proposals. Since 1 October 2026, both have been callable through Workers AI, and their BF16 weights, processor files, joint-head weights and reference Python implementation have been available under Apache 2.0. This article describes the published Clef revision 2f3de3d and Clef-flash revision 17f0b0a, checked on 2 October 2026.[^release][^clef-card][^flash-card]

They are not chat models. A request contains a state plus typed questions whose permitted answers are declared in advance. One forward pass returns one logit per permitted option; a separate softmax is applied within each question. There is no token-by-token answer generation and therefore no free-form output parser.[^clef-card][^implementation]

The two released variants

ComponentClefClef-flash
Post-trained backboneQwen3.8-27BQwen3.5-9B
Text stack64 layers, width 5,12032 layers, width 4,096
Sequence-mixing cadence48 linear-attention + 16 full-attention layers24 linear-attention + 8 full-attention layers
Full-attention layout24 query heads, four KV heads, head dimension 25616 query heads, four KV heads, head dimension 256
Joint schema headWidth 1,024; 16 heads; two routing layers; four field layers; FFN width 4,096Same, except for the 4,096-wide backbone input projection
Published BF16 parameters27.36B9.41B
Workers AI context65,536 tokens65,536 tokens

The layer and tensor values come from the released checkpoint configurations, not from rounded product names; the parameter totals are the repositories' safetensors metadata at the pinned revisions. Both Qwen backbones alternate three linear-attention layers with one full-attention layer and retain their vision encoder. This is related to, but not the same checkpoint as, Qwen3.8-Flash-Next. Clef does not add that model's n-gram memory.[^clef-card][^flash-card][^clef-config][^flash-config]

Input layout and the single backbone pass

The reference encoder turns a record into one causal sequence:

text
system instruction | state and optional media | schema fields and options | assistant suffix

For each field it records the token span of the question text and every option description. JSON states are serialized with sorted keys. Images and video frames are converted through the Qwen processor and inserted before the state. The schema is never silently shortened: if its fixed tokens exceed the configured limit, encoding fails. Otherwise, the state is truncated at the end so that the schema and suffix still fit.[^implementation]

Let LL be the valid input length, hh the backbone width, FF the number of fields and OfO_f the option count of field ff. After the Qwen prefill, the final hidden tensor is

H∈RLΓ—h,h∈{5120,4096}.H \in \mathbb{R}^{L \times h}, \qquad h \in \{5120,4096\}.

The release wrapper calls the text model with use_cache=False. It consumes the final hidden states immediately and does not retain a decode KV cache, because there is no autoregressive decode phase. This removes decode-time cache growth; it does not make the prefill free. Every fourth backbone layer still performs full causal attention, and the final HH plus a projected head memory both grow linearly with LL.[^implementation][^clef-config][^flash-config]

Stage 1: option-specific evidence routing

The joint head first layer-normalizes HH and projects every token to a 1,024-wide memory:

M=WmLN⁑(H)∈RLΓ—1024.M = W_m\operatorname{LN}(H) \in \mathbb{R}^{L \times 1024}.

For every question and option, the implementation mean-pools several spans. The resulting shapes for one record are:

TensorShapeSource
Question vectors QQ[F,h][F,h]Mean of the contextual question-token states
Option context CC[O,h][O,h], O=βˆ‘fOfO=\sum_f O_fMean of each contextual option span
Lexical option vectors EE[O,h][O,h]Mean of the backbone output-embedding rows for the option token IDs
Option queries XX[1,O,1024][1,O,1024]Sum of projected CC, EE, and the corresponding row of QQ

Two EvidenceRoutingLayers update XX. Each layer uses 16-head cross-attention from all option queries to MM, so its head dimension is 1024/16=641024/16=64. A residual 1,024β†’4,096β†’1,024 GELU feed-forward block follows. Options do not self-attend in this stage; each independently retrieves evidence from the complete encoded record. Its attention work is proportional to O LO\,L, rather than L2L^2, but more declared options still increase compute and temporary attention storage.[^head-config][^implementation]

Stage 2: joint fields and cross-field interaction

The routed option vectors are split back into their fields. For each field, a dot product between its projected question vector and its options produces a softmax over that field's options. The weighted option summary is added to four terms:

  1. the projected question vector;
  2. the option summary;
  3. a projection of the final sequence token;
  4. a learned embedding for the noul, choice, or score type.

This forms FF field vectors of width 1,024. Four standard pre-norm TransformerDecoderLayers then apply unmasked self-attention across the fields, cross-attention back to the full memory MM, and a 4,096-wide feed-forward block. The important distinction is concrete: stage 1 routes state evidence into individual options; stage 2 lets fields interact and revisit the state. Field self-attention costs O(F2)O(F^2), while the four memory cross-attentions cost O(F L)O(F\,L).[^implementation]

Schema-bound scoring and the lexical prior

Each final field vector is scored only against its own routed options. The implementation combines two paths:

+ \sigma(g)\left[s_j\,\cos(R_o,F_f)+\operatorname{MLP}(F_f,R_o,F_f\odot R_o,|F_f-R_o|)\right].$$ $R_o$ is the routed option vector, $F_f$ the final field vector, $E_o$ the lexical option vector, and $s_p$ and $s_j$ are learned positive scales clamped to at most 100. The residual MLP receives four 1,024-wide feature blocks. Finally, a softmax is applied separately to each field's variable-length logit vector.[^implementation] The lexical path is easy to misname. It looks up the existing Qwen output-embedding rows for the option's tokens and averages them. It is **not** an Engram-style hashed n-gram table, a vector database, or a new persistent memory. It is a per-request semantic prior over user-supplied answer labels. ## What was trained and what was released Cloudflare states that it froze the Qwen backbones and jointly trained the schema head with rank-256 low-rank adapters. The training objectives combined label-smoothed cross-entropy with a Brier loss, and a later RLCD objective gave partial credit to adjacent ordinal choices while penalizing drift from a reference policy. These are the vendor's training descriptions; the release does not include a full training recipe or datasets with which to independently reproduce the reported gains.[^release] The downloadable release is easier to characterize. `load_release_model` calls the checkpoint a **merged backbone**, loads it as ordinary `Qwen3_5ForConditionalGeneration` weights, and then attaches the separate `joint_head.safetensors`. There is no LoRA adapter file to load at inference. From the published head configuration, the head contains approximately 128.1 million parameters for Clef and 121.8 million for Clef-flash; those are derived parameter counts, not vendor-reported measurements.[^head-config][^implementation] ## Production boundaries Cloudflare's hosted endpoint is the production-ready path today: it accepts up to 64 questions, supports up to four embedded PNG, JPEG or WebP images under documented body and image limits, and exposes a 65,536-token context. The open-weight reference was tested by Cloudflare with PyTorch 2.11 and Transformers 5.10.2 on one H200. That is a reported test environment, not a claim that an H200 is the minimum hardware.[^workers][^clef-card] There is an important local/hosted mismatch. The released helper defaults `encode_record` and `systemone` to **16,384 tokens**, even though the checkpoint configuration allows 262,144 positions and Workers AI exposes 65,536. A local caller must raise `max_length` deliberately and then capacity-test the longer prefill. The custom head also processes records in a Python loop and returns nested variable-length tensors; that is a clear reference implementation, not evidence of optimized continuous batching in a general-purpose serving engine.[^implementation][^clef-config][^workers] Cloudflare publishes latency and quality tables, but those remain vendor-run measurements on its chosen suite and infrastructure. The architectural conclusion does not depend on them: Clef trades open-ended generation for one schema-constrained prefill plus a bounded decision head. Whether that is faster or better for a production workload depends on input length, number of questions and options, calibration requirements, GPU/runtime implementation, and the cost of using a generative fallback when the declared schema is incomplete. ## Sources [^release]: [Cloudflare, *Introducing Clef: our open-source decision models*, 1 October 2026](https://blog.cloudflare.com/clef-decision-models/). [^workers]: [Cloudflare Workers AI, official Clef model documentation and request limits](https://developers.cloudflare.com/workers-ai/models/clef/). [^clef-card]: [Cloudflare, Clef model card and released weights, revision `2f3de3d`](https://huggingface.co/Cloudflare/clef/tree/2f3de3dd85f379784083b0814d997ab627200f0c). [^flash-card]: [Cloudflare, Clef-flash model card and released weights, revision `17f0b0a`](https://huggingface.co/Cloudflare/clef-flash/tree/17f0b0ad64efb65d273590632833508766b2aae6). [^clef-config]: [Cloudflare, released Clef `config.json`](https://huggingface.co/Cloudflare/clef/blob/2f3de3dd85f379784083b0814d997ab627200f0c/config.json). [^flash-config]: [Cloudflare, released Clef-flash `config.json`](https://huggingface.co/Cloudflare/clef-flash/blob/17f0b0ad64efb65d273590632833508766b2aae6/config.json). [^head-config]: [Cloudflare, released joint-head configurations for Clef](https://huggingface.co/Cloudflare/clef/blob/2f3de3dd85f379784083b0814d997ab627200f0c/joint_head_config.json) [and Clef-flash](https://huggingface.co/Cloudflare/clef-flash/blob/17f0b0ad64efb65d273590632833508766b2aae6/joint_head_config.json). [^implementation]: [Cloudflare, released `joint_schema_model.py`, revision `2f3de3d`](https://huggingface.co/Cloudflare/clef/blob/2f3de3dd85f379784083b0814d997ab627200f0c/joint_schema_model.py).