Skip to content
Friendly disclaimer: flozi00 TechHub is a solo side-project next to a full-time job — personal learning notes, no official statements. Verify critical steps yourself.

Inside EmbeddingGemma 2: Soft Tokens, PLE, Hybrid Attention, and MRL

A source-checked guide to Google's released 740M multimodal embedding model: media soft tokens, per-layer embeddings, bidirectional local/global attention, mean pooling, and Matryoshka truncation.

8 min readflozi00
aiembeddingsarchitecturegemmamultimodalattentionretrievalinference

EmbeddingGemma 2 is a released open-weight embedding model, not a generative LLM or an architecture preview. Google DeepMind published the Apache-2.0 checkpoint on 6 October 2026. It maps text and code, images, sampled video frames, audio, or interleaved combinations into the same 768-dimensional vector space. The complete model has 740 million parameters, while deployments can omit the independent vision and audio encoders and retain a 270M text-only model.12

Its useful architectural idea is how ordinary encoder attention, modality-specific towers, and retrieval output are joined. Vision and audio towers do not produce final embeddings of their own. They produce soft tokens at the text model's width; those vectors replace media placeholders in one shared sequence. A 24-layer bidirectional text encoder then processes all modalities, keeps the original input directly accessible through Per-Layer Embeddings (PLE), and finally mean-pools token states into one normalized vector. This article describes the released checkpoint and pinned implementation. Benchmark and device-memory figures remain vendor measurements.

The released stack

ComponentReleased valueRuntime consequence
Text encoder24 layers, residual width 512, gated-GELU MLP width 2,048Every transformer block consumes and returns [B, S, 512].
Hybrid attention20 local and four global bidirectional layers in a repeating 5:1 patternLocal layers bound most attention work; four layers still form full S × S attention.
Local attentionFour query heads, two KV heads, head dimension 256Q is [B, 4, S, 256]; K and V are [B, 2, S, 256]. Each position sees a radius of 512 tokens.
Global attentionFour query heads, one KV head, head dimension 512Q is [B, 4, S, 512]; K and V are [B, 1, S, 512]. These layers use the complete unmasked input.
Position encodingRoPE base 10,000 locally and 1,000,000 globallyThe two attention paths rotate Q/K at different frequency scales.
Multimodal bridgeVision and audio towers project to 512-wide soft tokensMedia vectors enter the same residual stream as text embeddings.
OutputPer-token projection 512 -> 768, masked mean pooling, L2 normalizationOne input becomes one unit-length vector [B, 768].
Optional output sizesLeading 512, 256, or 128 dimensions through MRLShortening reduces vector-index storage and similarity work, not encoder compute.

The parameter count is deliberately modular: 130M parameters for the transformer backbone, 140M for embeddings and the output path, 170M for vision, and 300M for audio. The official configuration therefore supports text-only (270M), text plus vision (440M), text plus audio (570M), and complete multimodal (740M) loading. This is a weight-residency choice; it does not change the 512-wide fusion interface or the final 768-dimensional space.13

Media becomes soft tokens before the shared encoder

The processor builds a single token sequence and marks media positions with <|image|>, <|video|>, or <|audio|>. The corresponding modality tower emits a sequence of 512-wide vectors. In the published Transformers implementation, masked_scatter replaces the placeholder positions in inputs_embeds with those vectors. The resulting tensor has shape [B, S, 512] regardless of which modalities supplied its positions, so the same 24-layer encoder can mix text and media context.45

This is early fusion at the token level, not late fusion of separately pooled text, image, and audio embeddings. An interleaved product record, for example, can contain text, two image spans, and sampled video-frame spans; the model produces one vector for the complete sequence. The soft-token count is also the practical compute control:

  • an image uses 280 vision tokens by default;
  • a video frame uses 140 by default, sampled at 1 FPS and capped at 32 frames by the default processor;
  • audio uses about 25 tokens per second and should be supplied as mono 16 kHz;
  • the supported vision budgets are 70, 140, 280, 560, and 1,120 soft tokens.

All modalities share the released 8,192-token operating limit. With no accompanying text, the defaults fit roughly 29 images, 58 video frames, or 327 seconds of audio. Those are budget calculations, not simultaneous maxima: interleaved text and media consume the same sequence. A larger vision budget preserves more visual detail but increases local-attention work and, in the four global layers, the quadratic attention matrix.15

The checkpoint configuration contains a larger positional-capacity field, but Google documents and supports 8,192 tokens for this release. A configuration ceiling is not evidence for a larger validated production context.

Bidirectional hybrid attention has no reusable KV cache

All 24 layers are encoder layers: every valid token can attend left and right. Twenty local layers use a symmetric radius of 512, so an interior token can see at most 1,025 positions including itself. Every sixth layer is global, giving four full-attention layers. For sequence length S, the rough attention-score work is therefore

20×O(S×1025)+4×O(S2).20 \times O(S \times 1025) + 4 \times O(S^2).

This lowers the coefficient of quadratic attention but does not make the encoder linear-time. Media tokens count exactly like text positions after fusion.

The local layers use grouped-query attention: two KV heads are repeated across four query heads. The global layers use multi-query attention: one KV head is shared by four query heads. Queries, keys, and values receive per-head RMS normalization. The released eager implementation consequently uses an attention scale of 1.0 instead of applying another conventional 1 / sqrt(head_dim) factor, and it computes the softmax in FP32 before converting back to the value dtype.43

Unlike a causal decoder, this model does not append one token at a time and retain a cross-request KV cache. Embedding a changed document recomputes its complete sequence. GQA and MQA still reduce K/V projection width and transient tensors, but they do not create the persistent decode-cache benefit discussed in the KV-cache guide. Production batching should therefore be planned around encoder length and media-token budgets rather than generation concurrency.

PLE is a depth-wise input path, not an Engram lookup

EmbeddingGemma 2 uses the projection-only form of Per-Layer Embeddings. After media injection, the original [B, S, 512] input is projected once to [B, S, 24 × 512], scaled by 512^-0.5, reshaped to [B, S, 24, 512], and RMS-normalized. Layer i receives its own [B, S, 512] slice.4

After that layer's attention and MLP residual updates, the current hidden state passes through a learned 512 -> 512 projection and GELU. It is multiplied elementwise by the layer-specific PLE slice, projected again to 512, normalized, and added as another residual update:

hi′=hi+RMSNorm⁡(W2,i(GELU⁡(W1,ihi)⊙ei)).h_i' = h_i + \operatorname{RMSNorm}\left(W_{2,i}\left(\operatorname{GELU}(W_{1,i}h_i) \odot e_i\right)\right).

Here, e_i is derived from the input token vector at the same sequence position. The path lets every depth consult a learned transformation of the unmodified input instead of relying only on information carried through preceding residual blocks.

It is important not to call this an Engram memory. The released PLE path has no hash, n-gram key, external table, or sparse associative lookup. It is a dense, position-aligned projection of the current input sequence. The mechanism may serve a related high-level goal—preserving access to input identity—but its data structure and runtime behavior are different.

Mean pooling and Matryoshka truncation

The text encoder projects every final token state from 512 to 768 dimensions. Sentence Transformers then applies padding-mask-aware mean pooling and L2-normalizes the result in FP32. For valid-token mask m_t, the output before normalization is

z=∑tmtytmax⁡(∑tmt,ϵ),yt∈R768.z = \frac{\sum_t m_t y_t}{\max(\sum_t m_t, \epsilon)}, \qquad y_t \in \mathbb{R}^{768}.

Matryoshka Representation Learning makes useful information available in leading prefixes of this vector. A 256-dimensional index keeps z[:256], for example, and normalizes the sliced vector again. Slicing an already normalized 768d vector does not preserve unit length. Queries and corpus entries must also use the same dimension; otherwise their dot product is undefined.16

For FP32 storage, one raw vector occupies 3,072 bytes at 768d, 2,048 at 512d, 1,024 at 256d, or 512 at 128d, before index metadata. The 128d form is therefore six times smaller. It does not reduce transformer or modality-tower compute because the model still constructs and pools the 768d representation before truncation.

Google's full-precision benchmark table shows only small aggregate changes at 512d and 256d, but a larger decline at 128d—especially for the reported multimodal suite. These are vendor benchmark results, not a workload guarantee. The production decision is empirical: build the same recall and ranking evaluation at each candidate dimension, including the approximate-nearest-neighbor index settings.

Production boundaries

EmbeddingGemma 2 has a real deployment surface: a 1.49 GB safetensors checkpoint, Sentence Transformers and Transformers support, plus documented vLLM, SGLang, MLX, Ollama, LM Studio, and LiteRT paths. It is nevertheless an embedding component, not a complete search system. Chunking, metadata filters, ANN parameters, reranking, updates, and access control remain application responsibilities.27

Several boundaries matter more than the headline parameter count:

  1. Use BF16 or FP32, not FP16. Google states that the activation range exceeds FP16 and can yield NaNs or silently degraded vectors. BF16 retains FP32's exponent width; FP32 is the safe CPU fallback.1
  2. Use task prefixes consistently. Asymmetric retrieval requires different query and document prefixes. Prefixes apply to text, not image, video, or audio inputs.
  3. Treat 8K as a shared multimodal budget. A longer video or higher image-token budget directly displaces text and increases the four global attention costs.
  4. Load only used towers. Selective loading reduces resident weights. It does not make a 740M checkpoint dynamically skip a loaded tower without application control.
  5. Validate MRL per corpus. A smaller vector changes retrieval quality as well as storage and ANN behavior.

Google reports approximately 191 MB of active RAM for a quantized text-only deployment and 567 MB for the complete multimodal model on a Pixel 11 Pro. Those are vendor measurements for a particular device and quantized runtime, not sizes implied by the full-precision checkpoint and not universal RAM guarantees.7

The architecture's trade is concrete. Soft tokens provide one fusion interface but spend the same scarce context budget as text. Hybrid attention bounds most pairwise work but retains four quadratic global layers. PLE gives every layer a direct input path without an external memory table. MRL compresses the vector index but not the encoder pass. Those distinctions are what determine a production design; the 740M label alone does not.

Sources

Footnotes

  1. Google DeepMind, official EmbeddingGemma 2 model card, updated 6 October 2026. ↩ ↩2 ↩3 ↩4 ↩5

  2. Google, released embeddinggemma-2 checkpoint at pinned revision 914f7f8. ↩ ↩2

  3. Google, released embeddinggemma-2 config.json. ↩ ↩2

  4. Hugging Face, pinned Transformers implementation of EmbeddingGemma 2. ↩ ↩2 ↩3

  5. Hugging Face, pinned EmbeddingGemma 2 processing implementation. ↩ ↩2

  6. Google, released Sentence Transformers module configuration. ↩

  7. Google AI Edge, on-device EmbeddingGemma 2 deployment and measured Pixel 11 Pro memory. ↩ ↩2