Skip to content
Friendly disclaimer: flozi00 TechHub is a solo side-project next to a full-time job โ€” personal learning notes, no official statements. Verify critical steps yourself.

pplx-embed-v2: 9B documents, 0.6B queries

How Perplexity's aligned token embeddings enable asymmetric retrieval: MaxSim, ranking bounds, measured quality, indexing costs and deployment limits.

10 min readflozi00
aiembeddingsarchitectureretrievalragmath

The useful asymmetry

Documents usually change less often than people search them. A retrieval system can therefore spend more compute when building its index and less compute on each incoming query. Perplexity's released pplx-embed-v2-late-9b and pplx-embed-v2-late-0.6b support exactly this split: encode documents with the larger model, retain their token vectors, and encode queries with the smaller model. The October 6, 2026 release report describes this mixed configuration.

This article checks the public artifacts as of October 8, 2026. โ€œ9Bโ€ and โ€œ0.6Bโ€ below always mean the two late-interaction models. The separate pplx-embed-v2-context-9b-preview in the collection is outside this comparison. The late models have downloadable weights under MIT; the implementation is available through Sentence Transformers, rather than only through a hosted service.

What the encoders actually emit

These models do not collapse a document into one pooled vector. They emit contextual token vectors and compare those vectors after the two encoder forward passes. Text tokens, or visual tokens for an image or document page, become the retrievable representation. An image can thus be indexed directly; extracting readable passages for a later RAG answer remains a separate concern.

The 0.6B release configuration and 9B configuration describe Qwen3.5-based encoders with the following exported text-stack settings:

Configurationlate-0.6blate-9b
Text layers1232
Text hidden width1,0244,096
linear_attention layers624
full_attention layers68
Token projection1,024 โ†’ 1284,096 โ†’ 128
Projection biasfalsefalse
is_causalfalsefalse
query_length default1,0241,024
document_length default4,0964,096
query_expansionnullnull

The layer counts come from the actual layer_types arrays. In particular, the small model alternates six linear and six full-attention layers; inferring its schedule from the retained full_attention_interval: 4 field would give the wrong count. The sequence limits above come from each model's sentence_bert_config.json. A larger positional limit in the backbone configuration does not establish a supported retrieval length.

For a batch of text inputs, hidden states have shape [B,L,h][B,L,h]. A learned, bias-free projection maps the last dimension to 128. The exported module sequence is Transformer โ†’ Dense โ†’ MultiVectorMask โ†’ Normalize: masking removes unwanted positions, including configured document punctuation, and normalization acts on the retained token vectors. For one input, the effective result has shape [Lkept,128][L_{\mathrm{kept}},128]. There is no mean pooling at the end. See the 9B projection and Sentence Transformers module documentation.

For a retained hidden-state row zz, the output is:

u=zW,WโˆˆRhร—128,e=uโˆฅuโˆฅ2.u = zW,\qquad W\in\mathbb{R}^{h\times128},\qquad e = \frac{u}{\lVert u\rVert_2}.

The equation uses a row-vector convention; a PyTorch linear layer stores the transposed weight shape. The unit-norm expression assumes a nonzero projection. Actual normalization implementations handle very small norms numerically.

MaxSim: where the two models meet

Let QโˆˆRmร—128Q\in\mathbb{R}^{m\times128} contain the retained query vectors, and DโˆˆRnร—128D\in\mathbb{R}^{n\times128} the document vectors. The encoders can run at different times and on different machines. Scoring only needs their outputs.

A=QDโŠคโˆˆRmร—n,S(Q,D)=โˆ‘i=1mmaxโก1โ‰คjโ‰คnAij.A = QD^\top\in\mathbb{R}^{m\times n},\qquad S(Q,D)=\sum_{i=1}^{m}\max_{1\le j\le n} A_{ij}.

Each query token selects its best-matching document token; those maxima are summed. With normalized vectors, each dot product is a cosine similarity. Different query tokens may select the same document token: this is not a one-to-one assignment, and the score is not a probability. The official scoring API uses this sum by default, without query-length normalization.

For an illustrative two-token query about a battery warranty, suppose the similarities to three document tokens are:

A=[0.900.100.200.200.850.10],S=0.90+0.85=1.75.A=\begin{bmatrix} 0.90 & 0.10 & 0.20\\ 0.20 & 0.85 & 0.10 \end{bmatrix},\qquad S=0.90+0.85=1.75.

These numbers are invented to explain the reduction, not model measurements. Real token vectors are contextual and include the effects of tokenization and query/document formatting. Keeping local matches avoids requiring every aspect of a long document to survive in one pooled vector. That is a separate architectural benefit from mixing model sizes; compare the single-vector retrieval math and EmbeddingGemma 2 architecture.

Why equal dimensions are insufficient

Two encoders can both output 128 dimensions and still use incompatible coordinates. For an orthogonal matrix RR, rotating both sides preserves their dot product, while rotating only one side generally changes it:

(Rq)โŠค(Rd)=qโŠคd,(Rq)โŠคdโ‰ qโŠคd.(Rq)^\top(Rd)=q^\top d,\qquad (Rq)^\top d\ne q^\top d.

Consequently, good rankings within two separately trained embedding systems do not prove that their vectors can be mixed. Their coordinate systems must agree, including query/document conventions.

Perplexity reports that both released students were distilled from the same internal 18B teacher using per-token representation alignment, not ranking-distribution distillation alone. The relevant target is agreement with the teacher's coordinates. The LEAF paper, cited by the release, defines a representation loss using an unsquared Euclidean norm; Perplexity extends the idea to token representations. A schematic alignment objective is:

Lalign(ฮธ)=Ex[โˆ‘iโˆˆI(x)โˆฅfฮธ(x)iโˆ’T(x)iโˆฅ2].\mathcal{L}_{\mathrm{align}}(\theta) =\mathbb{E}_{x}\left[\sum_{i\in I(x)} \left\lVert f_\theta(x)_i-T(x)_i\right\rVert_2\right].

Here TT denotes the common teacher and I(x)I(x) the corresponding token positions used for alignment. This is an explanatory objective, not a reconstruction of the unpublished training recipe: loss weights, exact reduction, normalization during training and all token-correspondence details are not specified in the release. The architectural consequence is that the small query encoder and large document encoder approximate the same target representation space. The teacher is needed for training, not for serving the released pair.

How alignment controls scoring error

The following derivation explains the mechanism under explicit assumptions; it is not a measured error guarantee for these models. Assume normalized vectors and corresponding query/document token positions. Let qsq_s be a small-model query vector, dLd_L a large-model document vector, and qT,dTq_T,d_T their teacher references. Then:

qsโŠคdLโˆ’qTโŠคdT=(qsโˆ’qT)โŠคdL+qTโŠค(dLโˆ’dT).q_s^\top d_L-q_T^\top d_T =(q_s-q_T)^\top d_L+q_T^\top(d_L-d_T). โˆฃqsโŠคdLโˆ’qTโŠคdTโˆฃโ‰คโˆฅqsโˆ’qTโˆฅ2+โˆฅdLโˆ’dTโˆฅ2.\left|q_s^\top d_L-q_T^\top d_T\right| \le\lVert q_s-q_T\rVert_2+\lVert d_L-d_T\rVert_2.

The second line follows from Cauchyโ€“Schwarz and unit norms. Small coordinate errors imply small dot-product errors. MaxSim also behaves well when the winning document token changes, because the maximum is 1-Lipschitz in the largest elementwise error:

โˆฃmaxโกjajโˆ’maxโกjbjโˆฃโ‰คmaxโกjโˆฃajโˆ’bjโˆฃ.\left|\max_j a_j-\max_j b_j\right| \le\max_j|a_j-b_j|. ฮตi=โˆฅqs,iโˆ’qT,iโˆฅ2,ฮดj=โˆฅdL,jโˆ’dT,jโˆฅ2,\varepsilon_i=\lVert q_{s,i}-q_{T,i}\rVert_2,\qquad \delta_j=\lVert d_{L,j}-d_{T,j}\rVert_2, โˆฃS(Qs,DL)โˆ’S(QT,DT)โˆฃโ‰คโˆ‘i=1m(ฮตi+maxโกjฮดj).\left|S(Q_s,D_L)-S(Q_T,D_T)\right| \le\sum_{i=1}^{m}\left(\varepsilon_i+\max_j\delta_j\right).

A ranking requires another condition. If every candidate's score error is at most EE, a teacher score gap greater than 2E2E preserves the order of that candidate pair. With a smaller gap, either order is possible. A larger document encoder helps this bound if it reduces document representation error; parameter count alone does not prove that it does.

This explains why the mixed pair can retain useful rankings without matching the 9B query encoder exactly. It also explains the limits: small relevance margins can flip, query errors remain, and the public release does not provide the token-error values needed to turn this bound into a numerical guarantee. Comparing output vectors and retrieval rankings is therefore more informative than checking dimension equality alone.

What was actually measured

The release reports the following vendor measurements, not an independent reproduction. Higher nDCG@10 means better ranking under the evaluation's relevance labels; it is not a percentage of correctly answered queries.

Query / document encoderDomain-specific text nDCG@10ViDoRe v3 image nDCG@10
0.6B / 0.6B78.062.3
0.6B / 9Bโ‰ˆ79.6*63.5
9B / 9B81.365.2

The text evaluation averages six domain groups equally over 72 tasks. The mixed text value is derived from the rounded 78.0 baseline plus the reported 1.6-point improvement; it is not a separately published, more precise score. The mixed image value is stated directly. These results are from the release report's asymmetric-retrieval section.

From these rounded values, the mixed pair recovers approximately (79.6โˆ’78.0)/(81.3โˆ’78.0)โ‰ˆ48%(79.6-78.0)/(81.3-78.0)\approx48\% of the text-score gap and (63.5โˆ’62.3)/(65.2โˆ’62.3)โ‰ˆ41%(63.5-62.3)/(65.2-62.3)\approx41\% of the image-score gap. Those fractions describe these aggregate benchmarks. They do not promise the same gain on a particular corpus, and the mixed-model results do not establish per-query dominance. The section supplies no confidence intervals or query-serving latency measurements.

Where the compute savings come from

Let NN count document encoding events, including re-indexing, and UU count searches. Let csd,cLdc_s^d,c_L^d be the average document-encoding costs for the small and large models, and csq,cLqc_s^q,c_L^q their average query-encoding costs. For the same workload:

Cs/s=Ncsd+Ucsq,Cs/L=NcLd+Ucsq,CL/L=NcLd+UcLq.\begin{aligned} C_{s/s}&=Nc_s^d+Uc_s^q,\\ C_{s/L}&=Nc_L^d+Uc_s^q,\\ C_{L/L}&=Nc_L^d+Uc_L^q. \end{aligned} CL/Lโˆ’Cs/L=U(cLqโˆ’csq),Cs/Lโˆ’Cs/sU=NU(cLdโˆ’csd).C_{L/L}-C_{s/L}=U(c_L^q-c_s^q),\qquad \frac{C_{s/L}-C_{s/s}}{U} =\frac{N}{U}(c_L^d-c_s^d).

Compared with 9B throughout, the split saves the difference in query-encoding cost on every search. Compared with 0.6B throughout, it pays extra for document encoding; that extra cost is spread over UU searches. A mostly stable, frequently searched corpus is a natural fit. Frequently replaced documents or very few searches weaken the amortization.

These equations account for encoding only. They do not include the vector index, candidate retrieval, MaxSim, transfers or network latency. The nominal parameter ratio 9/0.6=159/0.6=15 is not a measured 15ร— speedup. Token lengths, attention schedule, batch size, precision, kernels and hardware all influence throughput and latency.

SetupPractical advantageCost or quality limit
0.6B throughoutLower document-encoding cost and a smaller query encoderLower reported retrieval scores
9B documents, 0.6B queriesHigher reported quality than 0.6B throughout; smaller online query encoder than 9B throughoutLarger document-encoding job; shared-space compatibility must be maintained
9B throughoutHighest reported scores among these three setupsLarger encoder in the query-serving path

The mixed setup also permits separate capacity planning: large-model indexing jobs can run on different hardware or schedules from the query service. That is a consequence of the independent encoder passes, not a published infrastructure benchmark.

The index and scorer remain substantial

The smaller query encoder does not automatically shrink stored document vectors. For equal retained token counts, dimensions and storage precision, all three setups have the same raw vector budget. For nn document vectors and bb bytes per scalar:

Mvectors=nโ‹…128โ‹…b.M_{\mathrm{vectors}}=n\cdot128\cdot b.

For 512 retained tokens, FP16 storage alone occupies 512โ‹…128โ‹…2=131,072512\cdot128\cdot2=131{,}072 bytes, or 128 KiB per document; FP32 doubles that. These are arithmetic examples of representation storage, not endorsed quantization settings or measured index sizes. Metadata, token IDs, index structures and replicas add overhead. Image-token counts depend on image preprocessing and resolution.

An exact score against one document computes an mร—nm\times n similarity matrix, with mnโ‹…128mn\cdot128 scalar multiplications and a similar number of additions. A corpus-scale system therefore needs a multi-vector candidate index and bounded candidate scoring; the library's dense similarity example is not such an index. Candidate pruning can lose relevant documents before MaxSim evaluates them. Vector quantization introduces another error source, discussed in quantization damage in retrieval.

The supported API and a deployment boundary

The model cards require sentence-transformers >= 6.0.0 and transformers >= 5.4.0. Use MultiVectorEncoder with encode_query and encode_document so the model's query/document prompts and token-level modules are applied. The native modules do not require trust_remote_code.

This small text example makes the two phases explicit. The first function belongs in the indexing job; its returned arrays are the document artifacts used by the query service. The second scores a supplied candidate set, not a full production corpus. The revisions pin the artifacts examined here.

python
from sentence_transformers import MultiVectorEncoder
 
DOC_MODEL = "perplexity-ai/pplx-embed-v2-late-9b"
DOC_REV = "77e936a1b18ed2ac00b7c76fccd70dc6a1bb1c18"
QUERY_MODEL = "perplexity-ai/pplx-embed-v2-late-0.6b"
QUERY_REV = "dd4e95b836a73f6f0c32e46ea127c0b86b02169e"
 
 
def build_document_vectors(texts):
    encoder = MultiVectorEncoder(DOC_MODEL, revision=DOC_REV)
    return encoder.encode_document(
        texts, batch_size=1, convert_to_numpy=True
    )
 
 
def score_candidates(queries, document_vectors):
    encoder = MultiVectorEncoder(QUERY_MODEL, revision=QUERY_REV)
    query_vectors = encoder.encode_query(
        queries, convert_to_numpy=True
    )
    return encoder.similarity(query_vectors, document_vectors)

In a running service, initialize and retain the query encoder once rather than constructing it for each request. Persist document arrays with their IDs and preprocessing/revision metadata; the example omits persistence and candidate indexing deliberately. Hardware placement and encoder precision must be chosen and evaluated for the actual workload.

The cards document separate text-only and image-only batches; mixed text-plus-image inputs are currently unsupported. Their formatting puts [Q] or [D] first, which differs from PyLate's conventional second-position marker. Do not substitute a generic pooled SentenceTransformer.encode call or assume another ColBERT pipeline is compatible because the output dimension is 128. The documented 1,024/4,096 sequence defaults also require a deliberate chunking or truncation policy for longer inputs.

Before choosing the mixed pair, compare all three configurations on the same corpus and relevance judgments. Keep preprocessing, candidate retrieval and scoring consistent; measure recall, nDCG, indexing throughput, query-encoding latency and end-to-end p50/p95 latency separately. Repeat the comparison after changing precision or quantizing vectors. Pin both model revisions and the encoder libraries, and check representation compatibility again after fine-tuning either side.

The supported reason to choose the split is specific: reuse the large model's document representations while serving queries with a smaller aligned encoder. Published measurements show a quality gain over the small model alone, with a remaining gap to the large model on both sides. Whether that tradeoff is worthwhile depends on corpus reuse, ranking margins, index costs and measured serving performance.

Primary sources and checked revisions

Sources were checked on October 8, 2026. The mathematical bounds and cost equations above are derived explanations; the benchmark values are attributed measurements, and the exported settings are artifact facts.