Clef and Clef-flash are released multimodal decision models from Cloudflare, not research-only architecture proposals. Since 1 October 2026, both have been callable through Workers AI, and their BF16 weights, processor files, joint-head weights and reference Python implementation have been available under Apache 2.0. This article describes the published Clef revision 2f3de3d and Clef-flash revision 17f0b0a, checked on 2 October 2026.[^release][^clef-card][^flash-card]
They are not chat models. A request contains a state plus typed questions whose permitted answers are declared in advance. One forward pass returns one logit per permitted option; a separate softmax is applied within each question. There is no token-by-token answer generation and therefore no free-form output parser.[^clef-card][^implementation]
The two released variants
| Component | Clef | Clef-flash |
|---|---|---|
| Post-trained backbone | Qwen3.8-27B | Qwen3.5-9B |
| Text stack | 64 layers, width 5,120 | 32 layers, width 4,096 |
| Sequence-mixing cadence | 48 linear-attention + 16 full-attention layers | 24 linear-attention + 8 full-attention layers |
| Full-attention layout | 24 query heads, four KV heads, head dimension 256 | 16 query heads, four KV heads, head dimension 256 |
| Joint schema head | Width 1,024; 16 heads; two routing layers; four field layers; FFN width 4,096 | Same, except for the 4,096-wide backbone input projection |
| Published BF16 parameters | 27.36B | 9.41B |
| Workers AI context | 65,536 tokens | 65,536 tokens |
The layer and tensor values come from the released checkpoint configurations, not from rounded product names; the parameter totals are the repositories' safetensors metadata at the pinned revisions. Both Qwen backbones alternate three linear-attention layers with one full-attention layer and retain their vision encoder. This is related to, but not the same checkpoint as, Qwen3.8-Flash-Next. Clef does not add that model's n-gram memory.[^clef-card][^flash-card][^clef-config][^flash-config]
Input layout and the single backbone pass
The reference encoder turns a record into one causal sequence:
system instruction | state and optional media | schema fields and options | assistant suffixFor each field it records the token span of the question text and every option description. JSON states are serialized with sorted keys. Images and video frames are converted through the Qwen processor and inserted before the state. The schema is never silently shortened: if its fixed tokens exceed the configured limit, encoding fails. Otherwise, the state is truncated at the end so that the schema and suffix still fit.[^implementation]
Let be the valid input length, the backbone width, the number of fields and the option count of field . After the Qwen prefill, the final hidden tensor is
The release wrapper calls the text model with use_cache=False. It consumes the final hidden states immediately and does not retain a decode KV cache, because there is no autoregressive decode phase. This removes decode-time cache growth; it does not make the prefill free. Every fourth backbone layer still performs full causal attention, and the final plus a projected head memory both grow linearly with .[^implementation][^clef-config][^flash-config]
Stage 1: option-specific evidence routing
The joint head first layer-normalizes and projects every token to a 1,024-wide memory:
For every question and option, the implementation mean-pools several spans. The resulting shapes for one record are:
| Tensor | Shape | Source |
|---|---|---|
| Question vectors | Mean of the contextual question-token states | |
| Option context | , | Mean of each contextual option span |
| Lexical option vectors | Mean of the backbone output-embedding rows for the option token IDs | |
| Option queries | Sum of projected , , and the corresponding row of |
Two EvidenceRoutingLayers update . Each layer uses 16-head cross-attention from all option queries to , so its head dimension is . A residual 1,024β4,096β1,024 GELU feed-forward block follows. Options do not self-attend in this stage; each independently retrieves evidence from the complete encoded record. Its attention work is proportional to , rather than , but more declared options still increase compute and temporary attention storage.[^head-config][^implementation]
Stage 2: joint fields and cross-field interaction
The routed option vectors are split back into their fields. For each field, a dot product between its projected question vector and its options produces a softmax over that field's options. The weighted option summary is added to four terms:
- the projected question vector;
- the option summary;
- a projection of the final sequence token;
- a learned embedding for the
noul,choice, orscoretype.
This forms field vectors of width 1,024. Four standard pre-norm TransformerDecoderLayers then apply unmasked self-attention across the fields, cross-attention back to the full memory , and a 4,096-wide feed-forward block. The important distinction is concrete: stage 1 routes state evidence into individual options; stage 2 lets fields interact and revisit the state. Field self-attention costs , while the four memory cross-attentions cost .[^implementation]
Schema-bound scoring and the lexical prior
Each final field vector is scored only against its own routed options. The implementation combines two paths:
+ \sigma(g)\left[s_j\,\cos(R_o,F_f)+\operatorname{MLP}(F_f,R_o,F_f\odot R_o,|F_f-R_o|)\right].$$ $R_o$ is the routed option vector, $F_f$ the final field vector, $E_o$ the lexical option vector, and $s_p$ and $s_j$ are learned positive scales clamped to at most 100. The residual MLP receives four 1,024-wide feature blocks. Finally, a softmax is applied separately to each field's variable-length logit vector.[^implementation] The lexical path is easy to misname. It looks up the existing Qwen output-embedding rows for the option's tokens and averages them. It is **not** an Engram-style hashed n-gram table, a vector database, or a new persistent memory. It is a per-request semantic prior over user-supplied answer labels. ## What was trained and what was released Cloudflare states that it froze the Qwen backbones and jointly trained the schema head with rank-256 low-rank adapters. The training objectives combined label-smoothed cross-entropy with a Brier loss, and a later RLCD objective gave partial credit to adjacent ordinal choices while penalizing drift from a reference policy. These are the vendor's training descriptions; the release does not include a full training recipe or datasets with which to independently reproduce the reported gains.[^release] The downloadable release is easier to characterize. `load_release_model` calls the checkpoint a **merged backbone**, loads it as ordinary `Qwen3_5ForConditionalGeneration` weights, and then attaches the separate `joint_head.safetensors`. There is no LoRA adapter file to load at inference. From the published head configuration, the head contains approximately 128.1 million parameters for Clef and 121.8 million for Clef-flash; those are derived parameter counts, not vendor-reported measurements.[^head-config][^implementation] ## Production boundaries Cloudflare's hosted endpoint is the production-ready path today: it accepts up to 64 questions, supports up to four embedded PNG, JPEG or WebP images under documented body and image limits, and exposes a 65,536-token context. The open-weight reference was tested by Cloudflare with PyTorch 2.11 and Transformers 5.10.2 on one H200. That is a reported test environment, not a claim that an H200 is the minimum hardware.[^workers][^clef-card] There is an important local/hosted mismatch. The released helper defaults `encode_record` and `systemone` to **16,384 tokens**, even though the checkpoint configuration allows 262,144 positions and Workers AI exposes 65,536. A local caller must raise `max_length` deliberately and then capacity-test the longer prefill. The custom head also processes records in a Python loop and returns nested variable-length tensors; that is a clear reference implementation, not evidence of optimized continuous batching in a general-purpose serving engine.[^implementation][^clef-config][^workers] Cloudflare publishes latency and quality tables, but those remain vendor-run measurements on its chosen suite and infrastructure. The architectural conclusion does not depend on them: Clef trades open-ended generation for one schema-constrained prefill plus a bounded decision head. Whether that is faster or better for a production workload depends on input length, number of questions and options, calibration requirements, GPU/runtime implementation, and the cost of using a generative fallback when the declared schema is incomplete. ## Sources [^release]: [Cloudflare, *Introducing Clef: our open-source decision models*, 1 October 2026](https://blog.cloudflare.com/clef-decision-models/). [^workers]: [Cloudflare Workers AI, official Clef model documentation and request limits](https://developers.cloudflare.com/workers-ai/models/clef/). [^clef-card]: [Cloudflare, Clef model card and released weights, revision `2f3de3d`](https://huggingface.co/Cloudflare/clef/tree/2f3de3dd85f379784083b0814d997ab627200f0c). [^flash-card]: [Cloudflare, Clef-flash model card and released weights, revision `17f0b0a`](https://huggingface.co/Cloudflare/clef-flash/tree/17f0b0ad64efb65d273590632833508766b2aae6). [^clef-config]: [Cloudflare, released Clef `config.json`](https://huggingface.co/Cloudflare/clef/blob/2f3de3dd85f379784083b0814d997ab627200f0c/config.json). [^flash-config]: [Cloudflare, released Clef-flash `config.json`](https://huggingface.co/Cloudflare/clef-flash/blob/17f0b0ad64efb65d273590632833508766b2aae6/config.json). [^head-config]: [Cloudflare, released joint-head configurations for Clef](https://huggingface.co/Cloudflare/clef/blob/2f3de3dd85f379784083b0814d997ab627200f0c/joint_head_config.json) [and Clef-flash](https://huggingface.co/Cloudflare/clef-flash/blob/17f0b0ad64efb65d273590632833508766b2aae6/joint_head_config.json). [^implementation]: [Cloudflare, released `joint_schema_model.py`, revision `2f3de3d`](https://huggingface.co/Cloudflare/clef/blob/2f3de3dd85f379784083b0814d997ab627200f0c/joint_schema_model.py).