Skip to content
Friendly disclaimer: flozi00 TechHub is a solo side-project next to a full-time job โ€” personal learning notes, no official statements. Verify critical steps yourself.

Clef Omni: Audio, Video, Sparse Experts and Schema Decisions

Inside Cloudflare's released Clef Omni: Qwen3-Omni encoders, synchronized audio and video, top-8 MoE routing, full grouped-query attention and a prefill-only schema head, with equations and deployment limits.

10 min readflozi00
aiarchitecturecloudflareqwenmultimodalaudiovideomoeattentioninference

Cloudflare released Clef Omni on 9 October 2026, with Apache 2.0 weights and a hosted Workers AI endpoint. It consumes a state, optional media and questions with explicitly allowed answers. Its output is a probability distribution over each question's options, rather than generated text. This article checks checkpoint revision 0db1cd2607d76a7bdb2a382f659e7b313079f84b against its configuration and reference code, using Transformers 5.10.2 as the implementation reference.1234

The existing Clef and Clef-flash guide explains the two-stage decision head. Omni changes the backbone: it uses the Qwen3-Omni-30B-A3B-Instruct Thinker, including audio and vision encoders, a sparse mixture of experts and full causal attention throughout the text stack. It does not use the dense variants' three-linear/one-full attention cadence.54

What actually runs

The loader sets enable_audio_output=False and retains the Thinker. The downloadable shards also contain the base model's speech-output modules, talker and code2wav, but those are excluded from this inference path. Input audio encoding remains enabled. Disabling speech output therefore does not turn Omni into a text-only model.236

ComponentReleased configuration
Thinker text stack48 layers; hidden width 2,048; vocabulary 152,064
Causal GQA per layer32 query heads; four KV heads; head dimension 128
Routed experts per layer128 experts; eight selected per token; expert intermediate width 768; no shared expert
Vision encoder27 blocks; width 1,152; output width 2,048; spatial patch 16; temporal patch 2; spatial merge 2ร—2
Audio encoder32 layers; width 1,280; 20 attention heads; 128 Mel bins; output width 2,048
Joint schema headWidth 1,024; 16 heads; two evidence-routing layers; four field layers; FFN width 4,096

These are checkpoint values, not specifications inferred from โ€œ30B-A3B.โ€ The text configuration allows 65,536 positions; both the released helper's default input limit and the hosted context are 64,000 tokens. Schema, media and text share that budget.5738

Media become embeddings, not transcripts

The reference encoder places media before the text state, followed by the complete schema and an assistant suffix. It records spans for question instructions and option descriptions. If media, schema and fixed framing already exceed the token budget, encoding fails; otherwise it truncates the end of the text state to fit. A long attachment can therefore remove text evidence without shortening the schema.3

The processor expands media placeholders into the appropriate number of token positions. The Thinker replaces the embeddings at those positions with encoder outputs of width 2,048. It does not generate an ASR transcript or image caption as an intermediate step. Both speech content and other acoustic features can enter this path, but that architectural ability does not establish accuracy for a particular sound or event.94

Audio: 16 kHz to roughly 13 positions per second

Encoded clips are decoded and resampled to 16 kHz mono. Raw sample arrays are taken as supplied, so callers must already provide that sample rate and layout. The feature extractor uses 128 Mel bins and a 160-sample hop: nominally 100 feature frames per second. Three stride-2 convolutions compress each 100-frame block to 13 positions; the audio encoder then projects its 1,280-wide representations to width 2,048.31094

For an unpadded Mel length M=100q+rM=100q+r, with 0โ‰คr<1000\le r<100, the published output-length function simplifies to the equation below. A 10-second clip with exactly 1,000 valid Mel frames contributes 130 audio positions, before framing tokens. Padding and boundary handling affect exact counts. This is a derived length example, not a measured throughput result.

Na(M)=13q+โŒˆr8โŒ‰,M=100q+r.N_a(M)=13q+\left\lceil\frac{r}{8}\right\rceil,\qquad M=100q+r.

Video: sampling, spatial merging and a common clock

The helper samples encoded video at two frames per second, scales frames to an approximate 262,144-pixel ceiling, and repeats the last frame when an even frame count is needed. Predecoded frame arrays must already follow the expected sampling cadence. The vision encoder groups pairs of frames, forms 16ร—16 spatial patches and merges 2ร—2 neighboring patches. For the processor's pre-merge grid (Tg,Hg,Wg)(T_g,H_g,W_g), the number of visual positions is:359

Nv=TgHgWg4.N_v=T_g\frac{H_gW_g}{4}.

A resized 384ร—384 frame gives Hg=Wg=24H_g=W_g=24, or 144 merged positions per temporal grid. At two frames per second and a temporal patch of two frames, that is approximately one such grid per second. This is a shape calculation; actual resizing, clip boundaries and marker tokens determine the request's final count.

For video with sound, the processor interleaves video and audio placeholders using a common temporal coordinate. At the default cadence, temporal grids are one second apart and their temporal positions advance by 13, matching approximately 13 audio positions per second. The implementation merges the streams in temporal order; it does not first transcribe audio and then attach text to frames.94

There is a concrete restriction: use_audio_in_video is enabled only when every video in the record has a decodable audio track. One silent video makes the helper omit all video soundtracks in that record. Separate entries in audio still work. The collator also rejects a batch mixing video records with and without this audio mode.3

Synchronization here concerns the sampled, encoded representations. Two-frame-per-second sampling can miss brief visual events, mono conversion removes separate audio channels, and the reference helper decodes the full clip before inference. This release path is not a streaming decision interface.

Position encoding and visual features

The Thinker uses interleaved multimodal RoPE. Its rotary position tensor has temporal, height and width coordinates. Of the 64 rotary frequencies for a 128-dimensional head, 24 use temporal coordinates, 20 height and 20 width; height and width frequencies are interleaved with temporal frequencies. The configured RoPE base is 10610^6. Text positions use the same index on all three axes, while visual positions retain their grid coordinates.54

Vision fusion also happens beyond the initial embedding replacement. Features from zero-based vision blocks 8, 16 and 24 are separately merged to width 2,048 and added at visual-token positions after text decoder layers 0, 1 and 2, respectively. These DeepStack additions reuse the visual positions; they do not append another three copies of the visual token sequence.54

Sparse experts reduce FFN work, not resident weights to 3B

For one token representation xโˆˆR2048x\in\mathbb{R}^{2048}, the router produces 128 logits, softmaxes them in FP32, selects the top eight probabilities and renormalizes only that subset. Each selected expert applies a SiLU-gated FFN with intermediate width 768. The routed output is a weighted sum:54

r=softmaxโก(Wrx),S=TopKโก8(r),ฮฑe=reโˆ‘jโˆˆSrj,y=โˆ‘eโˆˆSฮฑeWd,e[SiLUโก(Wg,ex)โŠ™(Wu,ex)].\begin{aligned} r&=\operatorname{softmax}(W_rx),\qquad S=\operatorname{TopK}_8(r),\\ \alpha_e&=\frac{r_e}{\sum_{j\in S}r_j},\\ y&=\sum_{e\in S}\alpha_e W_{d,e} \left[\operatorname{SiLU}(W_{g,e}x)\odot(W_{u,e}x)\right]. \end{aligned}

The fused gate/up weight tensor has shape [128,1536,2048][128,1536,2048]; the down weights have shape [128,2048,768][128,2048,768]. With no shared expert and a sparse block in every one of the 48 layers, the routed FFNs contain a derived 28,991,029,248 parameters. A token selects expert matrices containing 1,811,939,328 of those parameters across the stack. Attention, embeddings, routers and modality encoders are additional components; this is not a recalculation of the model name's total โ€œ3B activeโ€ figure.

All experts remain available to the router. The standard loader places the retained model on one device, rather than loading only eight experts. Cloudflare reports about 64 GB of BF16 backbone memory and testing on an H200 with PyTorch 2.11 and Transformers 5.10.2. That is the publisher's stated environment, not a measured minimum or a local measurement from this article.32

Relative to evaluating all 128 experts, selecting eight reduces expert arithmetic per token. It does not promise a 16ร— end-to-end speedup: full attention, encoders, dispatch, matrix utilization and memory traffic still cost time. Nor does the router assign fixed โ€œaudioโ€ or โ€œvideoโ€ roles to named experts; it routes each contextual token through learned scores.

Full GQA still has a prefill cost

For sequence length LL, each text layer projects [L,2048][L,2048] into Q:[32,L,128]Q:[32,L,128] and K,V:[4,L,128]K,V:[4,L,128], omitting the batch axis. Q/K undergo per-head RMS normalization and rotary encoding. Eight query heads share each KV head. The concatenated attention output is 4,096-wide and is projected back to 2,048; head width need not equal hidden width divided by query-head count.54

Every text layer uses full causal attention. Prefill still requires quadratically many token-pair comparisons, even though an IO-aware attention kernel can avoid materializing the complete score matrix. Transformers dispatches through its selected attention backend and permits alternative expert implementations; the Cloudflare loader does not pin one particular accelerated kernel. Kernel availability and an actual serving benchmark are separate from checkpoint architecture.43

Where Omni actually saves output work and memory

ClefModel saves the Thinker's lm_head weights for lexical option lookups and replaces the forward module with Identity. Consequently, the Thinker returns H:[L,2048]H:[L,2048] in its normally named logits field, rather than computing [L,152064][L,152064] vocabulary logits. It also runs with use_cache=False and does not call generate. The fixed <think>/</think> suffix is input framing, not a generated reasoning trace.3

The KV-cache formula gives a useful counterfactual: retaining BF16 K/V for all 48 layers would take 96 KiB per token, or 5.86 GiB at 64,000 positions. This cache is not retained by the release wrapper. The formula excludes allocator and runtime overhead.

BKV(L)=48โ‹…2โ‹…4โ‹…128โ‹…Lโ‹…2ย bytes.B_{\mathrm{KV}}(L)=48\cdot2\cdot4\cdot128\cdot L\cdot2\ \text{bytes}.

For another derived comparison, a BF16 [64000,152064][64000,152064] vocabulary tensor alone would occupy 18.13 GiB. The wrapper avoids that projection and tensor. Its final hidden tensor plus the head's [64000,1024][64000,1024] projected memory instead total 375 MiB at BF16. This is not peak VRAM: normalized copies, Q/K/V temporaries, attention and expert workspaces, encoder features and resident weights remain.

Skipping generation removes decode iterations and persistent decode-cache storage. It does not remove the multimodal encoders or the long full-attention prefill. The inference math guide explains why shorter output and smaller active FFNs cannot alone determine request latency.

From hidden states to bounded decisions

For one record with FF fields and O=โˆ‘fOfO=\sum_f O_f options, the head layer-normalizes HH and projects it to memory M:[1,L,1024]M:[1,L,1024]. It mean-pools the normalized question and option spans into 2,048-wide vectors, and averages the saved output-embedding rows for the option tokens. Their learned projections form option queries [1,O,1024][1,O,1024].73

Two 16-head cross-attention layers let each option gather evidence from the complete memory. Each head is 64-wide. The routed options are pooled within each field; their summaries join the projected question, final sequence state and a type embedding. Four decoder layers then apply unmasked field self-attention, memory cross-attention and an FFN to [1,F,1024][1,F,1024]. The backbone is causal, but this decision head can revisit the entire finished input.

Let eoe_o be the lexical option vector, qfq_f the pooled question and gg the final normalized sequence state. Let ufu_f and vov_o be the final normalized field and routed-option vectors. The head combines a 2,048-dimensional lexical prior with a 1,024-dimensional joint scorer:

โ„“o=spcosโก(eo,qf+g)+ฯƒ(ฮณ)[sjcosโก(uf,vo)+MLPโก([uf;vo;ufโŠ™vo;โˆฃufโˆ’voโˆฃ])].\ell_o=s_p\cos(e_o,q_f+g)+\sigma(\gamma)\left[ s_j\cos(u_f,v_o)+\operatorname{MLP}\bigl([u_f;v_o;u_f\odot v_o;|u_f-v_o|]\bigr) \right]. pf(o)=expโก(โ„“o)โˆ‘jโˆˆOfexpโก(โ„“j).p_f(o)=\frac{\exp(\ell_o)}{\sum_{j\in\mathcal{O}_f}\exp(\ell_j)}.

The positive learned scales sps_p and sjs_j are capped at 100, and the residual MLP receives 4,096 features. The lexical path uses ordinary vocabulary rows; it is not an Engram n-gram memory or external retrieval. Attention-pair counts in the head scale with 2OL+4(FL+F2)2OL+4(FL+F^2) per head, excluding projection and FFN work. Declaring more fields or options therefore still increases cost.3

The fields share representations, but the API returns separate distributions for each field, not a normalized distribution over all combinations of answers. Joint attention does not guarantee logically consistent field choices. A choice selects the largest probability; noul exposes p(true)p(\text{true}); score returns โˆ‘iiโ€‰p(i)\sum_i i\,p(i) over the zero-based ordered levels. That last output is an expected index, not a measured physical quantity.

Deployment boundaries and evidence

As checked on 10 October, Workers AI provides @cf/cloudflare/clef-omni with the following hosted limits. They are API restrictions, not limits enforced identically by the Python helper.8

Hosted inputLimit
Context / questions64,000 tokens / 1โ€“64 questions
ImagesFour; 4 MiB and 16 megapixels each; 8 MiB combined
Audio clipsFour; 8 MiB and 300 seconds each
VideosTwo; 16 MiB and 60 seconds each
Combined audio/video16 MiB decoded; embedded media, no remote URLs

Local code also accepts paths, URLs, bytes and predecoded arrays. It requires Pillow for images and PyAV for encoded audio/video. Pin the checkpoint, processor, custom code and runtime together: settings such as the audio sample rate or video cadence determine the data entering the network. The reference forward re-collates and runs each record sequentially without padding, even when passed a batch. It is not evidence of continuous batching or parallel batch execution.32

The pinned model card labels SGLang support as โ€œcoming soon.โ€ Cloudflare's accompanying announcement describes SGLang work for Clef; that does not demonstrate a released Omni serving integration. The hosted endpoint and the published local reference are the verified paths here.21

Cloudflare describes frozen-backbone post-training with LoRA and a combination of label-smoothed cross-entropy and Brier loss. Those objectives train option selection and probability estimates; they do not guarantee calibration on unseen inputs. The published quality and latency results are vendor measurements, and this article does not reproduce them. In particular, text decision benchmarks do not by themselves validate audio recognition or synchronized audiovisual decisions.12

Omni's concrete advantage over the dense Clef variants is a direct path for audio and audio-bearing video to the schema head. Its sparse FFNs reduce the amount of expert computation per token, while avoiding output generation removes a different cost. Whether those mechanisms improve a deployment depends on media quality, input length, schema size, calibration and the selected serving implementation. More modalities alone do not establish better decisions on text or images.

Primary sources

Footnotes

  1. Cloudflare, Clef Omni release announcement, 9 October 2026. โ†ฉ โ†ฉ2 โ†ฉ3

  2. Cloudflare, Clef Omni model card, revision 0db1cd2. โ†ฉ โ†ฉ2 โ†ฉ3 โ†ฉ4 โ†ฉ5 โ†ฉ6

  3. Cloudflare, released encoder, decision head and loader in joint_schema_model.py. โ†ฉ โ†ฉ2 โ†ฉ3 โ†ฉ4 โ†ฉ5 โ†ฉ6 โ†ฉ7 โ†ฉ8 โ†ฉ9 โ†ฉ10 โ†ฉ11 โ†ฉ12 โ†ฉ13

  4. Hugging Face Transformers 5.10.2, Qwen3-Omni model implementation, commit 0dad7b8. โ†ฉ โ†ฉ2 โ†ฉ3 โ†ฉ4 โ†ฉ5 โ†ฉ6 โ†ฉ7 โ†ฉ8 โ†ฉ9 โ†ฉ10

  5. Cloudflare, released config.json, revision 0db1cd2. โ†ฉ โ†ฉ2 โ†ฉ3 โ†ฉ4 โ†ฉ5 โ†ฉ6 โ†ฉ7

  6. Cloudflare, released weight index showing Thinker, Talker and Code2Wav tensors. โ†ฉ

  7. Cloudflare, released joint_head_config.json. โ†ฉ โ†ฉ2

  8. Cloudflare, official Clef Omni API documentation, checked 10 October 2026. โ†ฉ โ†ฉ2

  9. Hugging Face Transformers 5.10.2, Qwen3-Omni processor and audio/video interleaving. โ†ฉ โ†ฉ2 โ†ฉ3 โ†ฉ4

  10. Cloudflare, released processor_config.json. โ†ฉ