Cloudflare released Clef Omni on 9 October 2026, with Apache 2.0 weights and a hosted Workers AI endpoint. It consumes a state, optional media and questions with explicitly allowed answers. Its output is a probability distribution over each question's options, rather than generated text. This article checks checkpoint revision 0db1cd2607d76a7bdb2a382f659e7b313079f84b against its configuration and reference code, using Transformers 5.10.2 as the implementation reference.1234
The existing Clef and Clef-flash guide explains the two-stage decision head. Omni changes the backbone: it uses the Qwen3-Omni-30B-A3B-Instruct Thinker, including audio and vision encoders, a sparse mixture of experts and full causal attention throughout the text stack. It does not use the dense variants' three-linear/one-full attention cadence.54
What actually runs
The loader sets enable_audio_output=False and retains the Thinker. The downloadable shards also contain the base model's speech-output modules, talker and code2wav, but those are excluded from this inference path. Input audio encoding remains enabled. Disabling speech output therefore does not turn Omni into a text-only model.236
| Component | Released configuration |
|---|---|
| Thinker text stack | 48 layers; hidden width 2,048; vocabulary 152,064 |
| Causal GQA per layer | 32 query heads; four KV heads; head dimension 128 |
| Routed experts per layer | 128 experts; eight selected per token; expert intermediate width 768; no shared expert |
| Vision encoder | 27 blocks; width 1,152; output width 2,048; spatial patch 16; temporal patch 2; spatial merge 2ร2 |
| Audio encoder | 32 layers; width 1,280; 20 attention heads; 128 Mel bins; output width 2,048 |
| Joint schema head | Width 1,024; 16 heads; two evidence-routing layers; four field layers; FFN width 4,096 |
These are checkpoint values, not specifications inferred from โ30B-A3B.โ The text configuration allows 65,536 positions; both the released helper's default input limit and the hosted context are 64,000 tokens. Schema, media and text share that budget.5738
Media become embeddings, not transcripts
The reference encoder places media before the text state, followed by the complete schema and an assistant suffix. It records spans for question instructions and option descriptions. If media, schema and fixed framing already exceed the token budget, encoding fails; otherwise it truncates the end of the text state to fit. A long attachment can therefore remove text evidence without shortening the schema.3
The processor expands media placeholders into the appropriate number of token positions. The Thinker replaces the embeddings at those positions with encoder outputs of width 2,048. It does not generate an ASR transcript or image caption as an intermediate step. Both speech content and other acoustic features can enter this path, but that architectural ability does not establish accuracy for a particular sound or event.94
Audio: 16 kHz to roughly 13 positions per second
Encoded clips are decoded and resampled to 16 kHz mono. Raw sample arrays are taken as supplied, so callers must already provide that sample rate and layout. The feature extractor uses 128 Mel bins and a 160-sample hop: nominally 100 feature frames per second. Three stride-2 convolutions compress each 100-frame block to 13 positions; the audio encoder then projects its 1,280-wide representations to width 2,048.31094
For an unpadded Mel length , with , the published output-length function simplifies to the equation below. A 10-second clip with exactly 1,000 valid Mel frames contributes 130 audio positions, before framing tokens. Padding and boundary handling affect exact counts. This is a derived length example, not a measured throughput result.
Video: sampling, spatial merging and a common clock
The helper samples encoded video at two frames per second, scales frames to an approximate 262,144-pixel ceiling, and repeats the last frame when an even frame count is needed. Predecoded frame arrays must already follow the expected sampling cadence. The vision encoder groups pairs of frames, forms 16ร16 spatial patches and merges 2ร2 neighboring patches. For the processor's pre-merge grid , the number of visual positions is:359
A resized 384ร384 frame gives , or 144 merged positions per temporal grid. At two frames per second and a temporal patch of two frames, that is approximately one such grid per second. This is a shape calculation; actual resizing, clip boundaries and marker tokens determine the request's final count.
For video with sound, the processor interleaves video and audio placeholders using a common temporal coordinate. At the default cadence, temporal grids are one second apart and their temporal positions advance by 13, matching approximately 13 audio positions per second. The implementation merges the streams in temporal order; it does not first transcribe audio and then attach text to frames.94
There is a concrete restriction: use_audio_in_video is enabled only when every video in the record has a decodable audio track. One silent video makes the helper omit all video soundtracks in that record. Separate entries in audio still work. The collator also rejects a batch mixing video records with and without this audio mode.3
Synchronization here concerns the sampled, encoded representations. Two-frame-per-second sampling can miss brief visual events, mono conversion removes separate audio channels, and the reference helper decodes the full clip before inference. This release path is not a streaming decision interface.
Position encoding and visual features
The Thinker uses interleaved multimodal RoPE. Its rotary position tensor has temporal, height and width coordinates. Of the 64 rotary frequencies for a 128-dimensional head, 24 use temporal coordinates, 20 height and 20 width; height and width frequencies are interleaved with temporal frequencies. The configured RoPE base is . Text positions use the same index on all three axes, while visual positions retain their grid coordinates.54
Vision fusion also happens beyond the initial embedding replacement. Features from zero-based vision blocks 8, 16 and 24 are separately merged to width 2,048 and added at visual-token positions after text decoder layers 0, 1 and 2, respectively. These DeepStack additions reuse the visual positions; they do not append another three copies of the visual token sequence.54
Sparse experts reduce FFN work, not resident weights to 3B
For one token representation , the router produces 128 logits, softmaxes them in FP32, selects the top eight probabilities and renormalizes only that subset. Each selected expert applies a SiLU-gated FFN with intermediate width 768. The routed output is a weighted sum:54
The fused gate/up weight tensor has shape ; the down weights have shape . With no shared expert and a sparse block in every one of the 48 layers, the routed FFNs contain a derived 28,991,029,248 parameters. A token selects expert matrices containing 1,811,939,328 of those parameters across the stack. Attention, embeddings, routers and modality encoders are additional components; this is not a recalculation of the model name's total โ3B activeโ figure.
All experts remain available to the router. The standard loader places the retained model on one device, rather than loading only eight experts. Cloudflare reports about 64 GB of BF16 backbone memory and testing on an H200 with PyTorch 2.11 and Transformers 5.10.2. That is the publisher's stated environment, not a measured minimum or a local measurement from this article.32
Relative to evaluating all 128 experts, selecting eight reduces expert arithmetic per token. It does not promise a 16ร end-to-end speedup: full attention, encoders, dispatch, matrix utilization and memory traffic still cost time. Nor does the router assign fixed โaudioโ or โvideoโ roles to named experts; it routes each contextual token through learned scores.
Full GQA still has a prefill cost
For sequence length , each text layer projects into and , omitting the batch axis. Q/K undergo per-head RMS normalization and rotary encoding. Eight query heads share each KV head. The concatenated attention output is 4,096-wide and is projected back to 2,048; head width need not equal hidden width divided by query-head count.54
Every text layer uses full causal attention. Prefill still requires quadratically many token-pair comparisons, even though an IO-aware attention kernel can avoid materializing the complete score matrix. Transformers dispatches through its selected attention backend and permits alternative expert implementations; the Cloudflare loader does not pin one particular accelerated kernel. Kernel availability and an actual serving benchmark are separate from checkpoint architecture.43
Where Omni actually saves output work and memory
ClefModel saves the Thinker's lm_head weights for lexical option lookups and replaces the forward module with Identity. Consequently, the Thinker returns in its normally named logits field, rather than computing vocabulary logits. It also runs with use_cache=False and does not call generate. The fixed <think>/</think> suffix is input framing, not a generated reasoning trace.3
The KV-cache formula gives a useful counterfactual: retaining BF16 K/V for all 48 layers would take 96 KiB per token, or 5.86 GiB at 64,000 positions. This cache is not retained by the release wrapper. The formula excludes allocator and runtime overhead.
For another derived comparison, a BF16 vocabulary tensor alone would occupy 18.13 GiB. The wrapper avoids that projection and tensor. Its final hidden tensor plus the head's projected memory instead total 375 MiB at BF16. This is not peak VRAM: normalized copies, Q/K/V temporaries, attention and expert workspaces, encoder features and resident weights remain.
Skipping generation removes decode iterations and persistent decode-cache storage. It does not remove the multimodal encoders or the long full-attention prefill. The inference math guide explains why shorter output and smaller active FFNs cannot alone determine request latency.
From hidden states to bounded decisions
For one record with fields and options, the head layer-normalizes and projects it to memory . It mean-pools the normalized question and option spans into 2,048-wide vectors, and averages the saved output-embedding rows for the option tokens. Their learned projections form option queries .73
Two 16-head cross-attention layers let each option gather evidence from the complete memory. Each head is 64-wide. The routed options are pooled within each field; their summaries join the projected question, final sequence state and a type embedding. Four decoder layers then apply unmasked field self-attention, memory cross-attention and an FFN to . The backbone is causal, but this decision head can revisit the entire finished input.
Let be the lexical option vector, the pooled question and the final normalized sequence state. Let and be the final normalized field and routed-option vectors. The head combines a 2,048-dimensional lexical prior with a 1,024-dimensional joint scorer:
The positive learned scales and are capped at 100, and the residual MLP receives 4,096 features. The lexical path uses ordinary vocabulary rows; it is not an Engram n-gram memory or external retrieval. Attention-pair counts in the head scale with per head, excluding projection and FFN work. Declaring more fields or options therefore still increases cost.3
The fields share representations, but the API returns separate distributions for each field, not a normalized distribution over all combinations of answers. Joint attention does not guarantee logically consistent field choices. A choice selects the largest probability; noul exposes ; score returns over the zero-based ordered levels. That last output is an expected index, not a measured physical quantity.
Deployment boundaries and evidence
As checked on 10 October, Workers AI provides @cf/cloudflare/clef-omni with the following hosted limits. They are API restrictions, not limits enforced identically by the Python helper.8
| Hosted input | Limit |
|---|---|
| Context / questions | 64,000 tokens / 1โ64 questions |
| Images | Four; 4 MiB and 16 megapixels each; 8 MiB combined |
| Audio clips | Four; 8 MiB and 300 seconds each |
| Videos | Two; 16 MiB and 60 seconds each |
| Combined audio/video | 16 MiB decoded; embedded media, no remote URLs |
Local code also accepts paths, URLs, bytes and predecoded arrays. It requires Pillow for images and PyAV for encoded audio/video. Pin the checkpoint, processor, custom code and runtime together: settings such as the audio sample rate or video cadence determine the data entering the network. The reference forward re-collates and runs each record sequentially without padding, even when passed a batch. It is not evidence of continuous batching or parallel batch execution.32
The pinned model card labels SGLang support as โcoming soon.โ Cloudflare's accompanying announcement describes SGLang work for Clef; that does not demonstrate a released Omni serving integration. The hosted endpoint and the published local reference are the verified paths here.21
Cloudflare describes frozen-backbone post-training with LoRA and a combination of label-smoothed cross-entropy and Brier loss. Those objectives train option selection and probability estimates; they do not guarantee calibration on unseen inputs. The published quality and latency results are vendor measurements, and this article does not reproduce them. In particular, text decision benchmarks do not by themselves validate audio recognition or synchronized audiovisual decisions.12
Omni's concrete advantage over the dense Clef variants is a direct path for audio and audio-bearing video to the schema head. Its sparse FFNs reduce the amount of expert computation per token, while avoiding output generation removes a different cost. Whether those mechanisms improve a deployment depends on media quality, input length, schema size, calibration and the selected serving implementation. More modalities alone do not establish better decisions on text or images.
Primary sources
Footnotes
-
Cloudflare, Clef Omni release announcement, 9 October 2026. โฉ โฉ2 โฉ3
-
Cloudflare, Clef Omni model card, revision
0db1cd2. โฉ โฉ2 โฉ3 โฉ4 โฉ5 โฉ6 -
Cloudflare, released encoder, decision head and loader in
joint_schema_model.py. โฉ โฉ2 โฉ3 โฉ4 โฉ5 โฉ6 โฉ7 โฉ8 โฉ9 โฉ10 โฉ11 โฉ12 โฉ13 -
Hugging Face Transformers 5.10.2, Qwen3-Omni model implementation, commit
0dad7b8. โฉ โฉ2 โฉ3 โฉ4 โฉ5 โฉ6 โฉ7 โฉ8 โฉ9 โฉ10 -
Cloudflare, released
config.json, revision0db1cd2. โฉ โฉ2 โฉ3 โฉ4 โฉ5 โฉ6 โฉ7 -
Cloudflare, released weight index showing Thinker, Talker and Code2Wav tensors. โฉ
-
Cloudflare, official Clef Omni API documentation, checked 10 October 2026. โฉ โฉ2
-
Hugging Face Transformers 5.10.2, Qwen3-Omni processor and audio/video interleaving. โฉ โฉ2 โฉ3 โฉ4