Skip to content
Friendly disclaimer: flozi00 TechHub is a solo side-project next to a full-time job โ€” personal learning notes, no official statements. Verify critical steps yourself.

Mellum2.1: Sparse Experts, Local Attention and Selective YaRN

Inside JetBrains Mellum2.1: top-8 expert routing, explicit GQA tensor shapes, local and global attention, layer-selective YaRN and the difference between active parameters and serving memory.

10 min readflozi00
aiarchitecturejetbrainsmellummoeattentionropekv-cacheinference

JetBrains' Mellum2.1-12B-A2.5B-Thinking is an open-weight coding and reasoning model with Apache 2.0 licensing. Its October 2026 release updates post-training; JetBrains explicitly says the architecture is unchanged from Mellum2. This guide examines that previously released architecture through the actual 2.1 checkpoint, rather than presenting reinforcement learning as a new attention mechanism.12

The interesting combination is top-8-of-64 expert routing, four KV heads and three local-attention layers for every global layer. Long-context position scaling applies only to the global layers. These choices reduce different parts of inference cost, and none makes the entire model a 2.5-billion-parameter allocation.

The artifacts being explained

Checked on 11 October 2026, the BF16 checkpoint is revision 92ddae9fc7665e9f801d141d2e5a6b2caf2460c4, last modified on 7 October. The configuration declares MellumForCausalLM, not a generic Qwen class. Implementation references below are Transformers 5.10.2, commit 0dad7b822255a0ae261ec45ae937371e859ffd1a, and vLLM 0.31.0, commit db9527a46873454610df6dbedf79a36d6bf1a7f6.3456

The official GGUF repository already contains BF16, Q8_0, Q6_K, Q4_K_M and MXFP4_MOE files at revision 20b439426f1e997ba575eaccaf4856ac0f76f5d5. That artifact inventory is more current than the BF16 card and launch post, which still describe GGUF builds as forthcoming. The MTP head remains announced separately: the inspected BF16 weight index has no MTP tensors. Ordinary inference with this checkpoint does not include the advertised speculative-decoding speedup.728

ComponentCheckpoint value
Decoder28 layers; residual width 2,304; RMSNorm epsilon 10โปโถ
Attention32 query heads; four KV heads; head width 128; QK normalization
Local/global pattern21 sliding layers, window 1,024; seven full-attention layers
Feed-forward64 routed experts per layer; eight selected; intermediate width 896; no shared expert
Vocabulary98,304; separate input embedding and output projection
Position limit131,072; global YaRN factor 16 from an 8,192-token reference

These values come from the pinned configuration and implementation. Every released layer has mlp_layer_types="sparse"; the intermediate_size=7168 field describes a possible dense MLP, not an additional dense branch in this checkpoint.45

For the ordinary token-to-output pipeline, start with how LLMs work. Mellum retains causal next-token generation; its differences lie inside attention and the feed-forward stage.

A block, with the dimensions made explicit

Let the incoming residual state be XโˆˆRBร—Tร—2304X\in\mathbb{R}^{B\times T\times2304} for batch size BB and the TT tokens processed in this call. A block first applies RMSNorm, attention and a residual addition, then a second RMSNorm, the sparse feed-forward network and another residual addition. Query and key vectors also receive RMSNorm across their 128 features per head, before rotary position encoding.5

The query projection is wider than the residual state. Head width is explicitly configured as 128; calculating it as 2304/32=722304/32=72 would give the wrong weights, rotary frequencies and cache budget.

TensorShape before attention
Q projection[B, T, 4096] โ†’ [B, 32, T, 128]
K projection[B, T, 512] โ†’ [B, 4, T, 128]
V projection[B, T, 512] โ†’ [B, 4, T, 128]
Attention result[B, T, 4096] โ†’ output projection โ†’ [B, T, 2304]

With grouped-query attention, each group of eight query heads uses one KV head. Each query still forms its own attention distribution; keys and values are shared, not the query results. The eager reference expands KV heads for the calculation, whereas an optimized GQA kernel can reuse them directly. The persistent cache needs only four heads.5

For one query head and position tt, attention has the usual form below. g(h)g(h) selects its KV group, and MtjM_{tj} is zero for visible positions and negative infinity elsewhere. Rotary-transformed Q and K are denoted with tildes.

ot,h=โˆ‘jsoftmaxโกjโ€‰โฃ(q~t,hโŠคk~j,g(h)128+Mtj)vj,g(h).o_{t,h}=\sum_j\operatorname{softmax}_j\!\left(\frac{\widetilde q_{t,h}^{\top}\widetilde k_{j,g(h)}}{\sqrt{128}}+M_{tj}\right)v_{j,g(h)}.

Local work, with periodic global access

Zero-based layers 3, 7, 11, 15, 19, 23 and 27 use full causal attention. The other 21 layers use a sliding causal window. In the Transformers reference, a local query at position tt can see keys satisfying tโˆ’1024<jโ‰คtt-1024<j\le t: at most 1,024 positions, including itself.49

For example, at position 20,000 a local layer directly sees positions 18,977 through 20,000. The following global layer can directly retrieve from the entire preceding sequence. Local layers can also carry information already incorporated into nearby hidden states; their window is not a rule that erases all influence of older text. Nevertheless, unrestricted retrieval is available in only seven layers, so this is a different learned computation from a 28-layer full-attention model.

For a long prefill of length LL, count permitted query-key pairs per head, ignoring padding. A full causal layer has L(L+1)/2L(L+1)/2 pairs. A local layer has โˆ‘t=1Lminโก(t,1024)\sum_{t=1}^{L}\min(t,1024), approximately 1024L1024L when LL is much larger than the window. The hybrid stack therefore has the following asymptotic pair count:

Phybridโ‰ˆ7L22+21โ‹…1024L,Pallย fullโ‰ˆ28L22.P_{\mathrm{hybrid}}\approx7\frac{L^2}{2}+21\cdot1024L, \qquad P_{\mathrm{all\ full}}\approx28\frac{L^2}{2}.

This is an operation-count derivation, not a runtime measurement. It reduces attention work only when the selected kernel exploits the local window. Projection matrices, experts, scheduling and memory transfers still cost time; an eager path that materializes dense masked scores need not realize the arithmetic saving.

Why only global layers use YaRN

RoPE rotates pairs of Q/K features as a function of token position. For 64 feature pairs and base ฮธ=500000\theta=500000, ordinary inverse frequencies are ฯ‰i=ฮธโˆ’2i/128\omega_i=\theta^{-2i/128}, for i=0,โ€ฆ,63i=0,\ldots,63. The same orthogonal rotations on Q and K make their positional interaction depend on relative distance:

(R(t)q)โŠค(R(j)k)=qโŠคR(jโˆ’t)k.(R(t)q)^{\top}(R(j)k)=q^{\top}R(j-t)k.

This identity explains the local/global distinction. A local layer sees relative distances no larger than 1,023 even late in a long sequence. Extending absolute token indices does not itself require a new distribution of relative distances for that layer. A global layer must compare positions separated by up to 131,071 tokens. Mellum's long-context training therefore keeps ordinary RoPE in local layers and applies YaRN only in global layers.104

Uniform position interpolation would divide every inverse frequency by 16. That stretches the usable position range but also changes short-distance rotations in every feature pair. YaRN instead blends unscaled and interpolated frequencies: fast-rotating pairs retain their original frequencies, slow pairs are stretched, and a transition band joins them.11

ri=clipโกโ€‰โฃ(iโˆ’1835โˆ’18,0,1),ฯ‰iglobal=(1โˆ’ri)ฯ‰i+riฯ‰i16,ฯ‰ilocal=ฯ‰i.\begin{aligned} r_i&=\operatorname{clip}\!\left(\frac{i-18}{35-18},0,1\right),\\ \omega_i^{\mathrm{global}}&=(1-r_i)\omega_i+r_i\frac{\omega_i}{16},\\ \omega_i^{\mathrm{local}}&=\omega_i. \end{aligned}

The bounds 18 and 35 are derived from the implementation's floor/ceil correction range with original length 8,192 and beta_fast=32, beta_slow=1. It locates feature-pair indices by how many rotations they complete over the original context: i(ฮฒ)=64lnโก(8192/(2ฯ€ฮฒ))/lnโก(500000)i(\beta)=64\ln(8192/(2\pi\beta))/\ln(500000). At indices 0โ€“18 the global frequency stays unchanged; at 35โ€“63 it is divided by 16. The intervening pairs are blended.11

The checkpoint also specifies an attention factor a=1.2772588722239782=1+0.1lnโก16a=1.2772588722239782=1+0.1\ln16. The implementation multiplies both sine and cosine by aa in global layers. Both rotated Q and K are consequently scaled, so the pre-softmax dot product acquires a2โ‰ˆ1.63139a^2\approx1.63139. Local layers use factor one. Reproducing only the frequency ramp while omitting this amplitude factor changes inference.4511

The report's rationale and the rotation identity support preserving local geometry. They do not prove that every long-context task succeeds or that this recipe dominates all alternatives. The original report's YaRN comparison also documents a QA prompt-formatting issue in its RULER evaluation. Those extension experiments concern Mellum2 training, not a fresh measurement of the released 2.1 checkpoint.10

Eight experts are computation, not an eight-expert model

For each normalized token vector xโˆˆR2304x\in\mathbb{R}^{2304}, the router projects to 64 logits, computes softmax in float32, selects the eight largest probabilities and renormalizes the selected set. Each selected expert uses a SiLU-gated MLP with intermediate width 896. Its output is weighted and accumulated into the token's 2,304-dimensional result.5

p=softmaxโก(Wrx),S=Top8โก(p),p^e=peโˆ‘uโˆˆSpu,fe(x)=Wedown[SiLUโก(Wegatex)โŠ™(Weupx)],y=โˆ‘eโˆˆSp^efe(x).\begin{aligned} p&=\operatorname{softmax}(W_rx),\quad S=\operatorname{Top8}(p),\quad \widehat p_e=\frac{p_e}{\sum_{u\in S}p_u},\\ f_e(x)&=W_e^{\mathrm{down}}\bigl[\operatorname{SiLU}(W_e^{\mathrm{gate}}x)\odot(W_e^{\mathrm{up}}x)\bigr],\\ y&=\sum_{e\in S}\widehat p_e f_e(x). \end{aligned}

Each expert has three matrices containing 3โ‹…2304โ‹…8963\cdot2304\cdot896 weights. Across 28 layers, all 64 experts contain about 11.10 billion weights, while the eight selected experts contribute about 1.39 billion weights to one token's path. Attention, routers, embeddings and the output projection are additional. These are derived matrix counts; โ€œ2.5B activeโ€ is a rounded model descriptor, not a guaranteed dense-model speed equivalence.

The published tensor metadata reports 12,149,923,072 BF16 parameters overall, or about 22.63 GiB at two bytes each, before cache and runtime memory. Unselected experts remain stored unless a deployment explicitly offloads or partitions them. Different tokens in a batch can collectively activate many more than eight experts. vLLM routes this model through its fused MoE implementation; the reference Python expert loop describes semantics, not the GPU kernel's execution schedule.312

A concrete hybrid KV-cache budget

For one request, one layer and one retained position, BF16 K plus V require 2โ‹…4โ‹…128โ‹…2=20482\cdot4\cdot128\cdot2=2048 bytes. GQA already reduces this by eight relative to 32 separately cached KV heads. Sliding attention then limits the positions needed in 21 layers; the seven global layers continue to grow with sequence length.4

Using a 1,024-position local working window, the logical K/V payload is below. This assumes an unsharded model, BF16 cache, no prefix sharing and a runtime that retains only the required local history. It includes the current position and excludes allocator overhead.

MKV(L)=2048[7L+21minโก(L,1024)]ย bytes.M_{\mathrm{KV}}(L)=2048\left[7L+21\min(L,1024)\right]\ \text{bytes}.
Sequence lengthAll 28 layers full, GQAHybrid local/global, GQA
8,1920.4375 GiB0.15039 GiB
32,7681.7500 GiB0.47852 GiB
131,0727.0000 GiB1.79102 GiB

At the configured maximum, the hybrid payload is 74.41% smaller than the hypothetical all-full version with the same heads. This is a computed storage comparison, not a measured VRAM reduction or a comparison with a separately trained model. The seven global caches still hold 1.75 GiB; local layers add about 0.041 GiB. Multiplying by the number of independent requests estimates their logical payload, not total server VRAM.

The Transformers dynamic cache keeps the last 1,023 local positions for the next call and combines them with new tokens for attention. Its update returns the larger concatenated tensors, and its stored slices can share their backing allocation. In particular, a large prefill or chunk can have a larger physical footprint than the retained tensor shape suggests. Paged allocation, temporary buffers and static caches can add further space. A local attention mask alone does not establish eviction or actual memory usage.13

Even the idealized maximum-context budget is 22.63 + 1.79 โ‰ˆ 24.42 GiB for BF16 weights and KV alone. Activations, scratch space and serving reservations come on top. This explains why a โ€œ2.5B activeโ€ model is not automatically comfortable in a 24 GiB budget. For allocation and concurrency details, see KV-cache memory math and quantization.

Serving the released path

The pinned vLLM implementation selects rope_parameters[layer_type] and passes the per-layer window into attention. It also handles the explicit 128-wide heads. A runtime that merely loads expert weights but applies one global RoPE configuration to all layers does not reproduce this checkpoint.6

With vLLM 0.31.0 already installed, the following starts the pinned BF16 model at a deliberately smaller 32,768-token limit. The parser flags follow the model card; they separate reasoning output and parse tool calls, rather than executing tools. Raise the context limit only after validating the memory budget and workload.2

bash
vllm serve JetBrains/Mellum2.1-12B-A2.5B-Thinking \
  --revision 92ddae9fc7665e9f801d141d2e5a6b2caf2460c4 \
  --dtype bfloat16 \
  --max-model-len 32768 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser hermes

This command is a source-checked configuration example, not an independently executed GPU test. Verify the selected attention backend supports the mixed window pattern and examine actual cache reservations. Tensor parallelism above four ranks can replicate KV heads in the vLLM implementation, so aggregate memory does not follow a naive division by rank count.6

GGUF lowers weight storage separately from the attention design. The official Q4_K_M file is listed as 8.1 GB; that is a file size, not total working memory. Its reported quantization comparisons use Wikitext-2 with a 512-token context. Those logit comparisons do not establish agentic coding accuracy or quality at 128K, and quantized weights do not automatically imply a quantized KV cache.7

Mellum2.1 provides a concrete example of composing sparse feed-forward computation with cheaper local attention while retaining periodic global access. Layer-selective YaRN adapts the global path's positional geometry. Whether that balance suits a coding service depends on its tasks, request lengths, quantization and runtime; architecture and parameter counts alone cannot establish throughput or answer quality.

Primary sources

Footnotes

  1. JetBrains, Mellum2.1 Gets to Work, October 2026; checked 2026-10-11. Explicitly distinguishes post-training changes from unchanged architecture. โ†ฉ

  2. JetBrains, Mellum2.1 Thinking model card, pinned revision. โ†ฉ โ†ฉ2 โ†ฉ3

  3. Hugging Face, JetBrains checkpoint metadata at the inspected revision, BF16 tensor count; checked 2026-10-11. โ†ฉ โ†ฉ2

  4. JetBrains, released configuration. โ†ฉ โ†ฉ2 โ†ฉ3 โ†ฉ4 โ†ฉ5 โ†ฉ6

  5. Hugging Face Transformers 5.10.2, Mellum implementation: MellumAttention, MellumRotaryEmbedding, MellumTopKRouter, MellumExperts, MellumDecoderLayer. โ†ฉ โ†ฉ2 โ†ฉ3 โ†ฉ4 โ†ฉ5 โ†ฉ6

  6. vLLM 0.31.0, Mellum implementation, per-layer rotary/window configuration and tensor-parallel KV-head allocation. โ†ฉ โ†ฉ2 โ†ฉ3

  7. JetBrains, official GGUF model card and files, pinned revision; card lists sizes and quantization measurement conditions. โ†ฉ โ†ฉ2

  8. JetBrains, released weight index. โ†ฉ

  9. Transformers 5.10.2, mask implementation, sliding_window_overlay, sliding_window_causal_mask_function. โ†ฉ

  10. Kojic et al., Mellum 2 Technical Report, 2026-05-29, especially Sections 2.2, 4.1 and C.1. Describes the inherited architecture and limits of its original extension evaluation. โ†ฉ โ†ฉ2

  11. Transformers 5.10.2, YaRN implementation, _compute_yarn_parameters. โ†ฉ โ†ฉ2 โ†ฉ3

  12. vLLM 0.31.0, Qwen3 MoE implementation reused by Mellum, Qwen3MoeSparseMoeBlock and FusedMoEFactory. โ†ฉ

  13. Transformers 5.10.2, cache implementation, DynamicSlidingWindowLayer.update and DynamicCache. โ†ฉ