JetBrains' Mellum2.1-12B-A2.5B-Thinking is an open-weight coding and reasoning model with Apache 2.0 licensing. Its October 2026 release updates post-training; JetBrains explicitly says the architecture is unchanged from Mellum2. This guide examines that previously released architecture through the actual 2.1 checkpoint, rather than presenting reinforcement learning as a new attention mechanism.12
The interesting combination is top-8-of-64 expert routing, four KV heads and three local-attention layers for every global layer. Long-context position scaling applies only to the global layers. These choices reduce different parts of inference cost, and none makes the entire model a 2.5-billion-parameter allocation.
The artifacts being explained
Checked on 11 October 2026, the BF16 checkpoint is revision 92ddae9fc7665e9f801d141d2e5a6b2caf2460c4, last modified on 7 October. The configuration declares MellumForCausalLM, not a generic Qwen class. Implementation references below are Transformers 5.10.2, commit 0dad7b822255a0ae261ec45ae937371e859ffd1a, and vLLM 0.31.0, commit db9527a46873454610df6dbedf79a36d6bf1a7f6.3456
The official GGUF repository already contains BF16, Q8_0, Q6_K, Q4_K_M and MXFP4_MOE files at revision 20b439426f1e997ba575eaccaf4856ac0f76f5d5. That artifact inventory is more current than the BF16 card and launch post, which still describe GGUF builds as forthcoming. The MTP head remains announced separately: the inspected BF16 weight index has no MTP tensors. Ordinary inference with this checkpoint does not include the advertised speculative-decoding speedup.728
| Component | Checkpoint value |
|---|---|
| Decoder | 28 layers; residual width 2,304; RMSNorm epsilon 10โปโถ |
| Attention | 32 query heads; four KV heads; head width 128; QK normalization |
| Local/global pattern | 21 sliding layers, window 1,024; seven full-attention layers |
| Feed-forward | 64 routed experts per layer; eight selected; intermediate width 896; no shared expert |
| Vocabulary | 98,304; separate input embedding and output projection |
| Position limit | 131,072; global YaRN factor 16 from an 8,192-token reference |
These values come from the pinned configuration and implementation. Every released layer has mlp_layer_types="sparse"; the intermediate_size=7168 field describes a possible dense MLP, not an additional dense branch in this checkpoint.45
For the ordinary token-to-output pipeline, start with how LLMs work. Mellum retains causal next-token generation; its differences lie inside attention and the feed-forward stage.
A block, with the dimensions made explicit
Let the incoming residual state be for batch size and the tokens processed in this call. A block first applies RMSNorm, attention and a residual addition, then a second RMSNorm, the sparse feed-forward network and another residual addition. Query and key vectors also receive RMSNorm across their 128 features per head, before rotary position encoding.5
The query projection is wider than the residual state. Head width is explicitly configured as 128; calculating it as would give the wrong weights, rotary frequencies and cache budget.
| Tensor | Shape before attention |
|---|---|
| Q projection | [B, T, 4096] โ [B, 32, T, 128] |
| K projection | [B, T, 512] โ [B, 4, T, 128] |
| V projection | [B, T, 512] โ [B, 4, T, 128] |
| Attention result | [B, T, 4096] โ output projection โ [B, T, 2304] |
With grouped-query attention, each group of eight query heads uses one KV head. Each query still forms its own attention distribution; keys and values are shared, not the query results. The eager reference expands KV heads for the calculation, whereas an optimized GQA kernel can reuse them directly. The persistent cache needs only four heads.5
For one query head and position , attention has the usual form below. selects its KV group, and is zero for visible positions and negative infinity elsewhere. Rotary-transformed Q and K are denoted with tildes.
Local work, with periodic global access
Zero-based layers 3, 7, 11, 15, 19, 23 and 27 use full causal attention. The other 21 layers use a sliding causal window. In the Transformers reference, a local query at position can see keys satisfying : at most 1,024 positions, including itself.49
For example, at position 20,000 a local layer directly sees positions 18,977 through 20,000. The following global layer can directly retrieve from the entire preceding sequence. Local layers can also carry information already incorporated into nearby hidden states; their window is not a rule that erases all influence of older text. Nevertheless, unrestricted retrieval is available in only seven layers, so this is a different learned computation from a 28-layer full-attention model.
For a long prefill of length , count permitted query-key pairs per head, ignoring padding. A full causal layer has pairs. A local layer has , approximately when is much larger than the window. The hybrid stack therefore has the following asymptotic pair count:
This is an operation-count derivation, not a runtime measurement. It reduces attention work only when the selected kernel exploits the local window. Projection matrices, experts, scheduling and memory transfers still cost time; an eager path that materializes dense masked scores need not realize the arithmetic saving.
Why only global layers use YaRN
RoPE rotates pairs of Q/K features as a function of token position. For 64 feature pairs and base , ordinary inverse frequencies are , for . The same orthogonal rotations on Q and K make their positional interaction depend on relative distance:
This identity explains the local/global distinction. A local layer sees relative distances no larger than 1,023 even late in a long sequence. Extending absolute token indices does not itself require a new distribution of relative distances for that layer. A global layer must compare positions separated by up to 131,071 tokens. Mellum's long-context training therefore keeps ordinary RoPE in local layers and applies YaRN only in global layers.104
Uniform position interpolation would divide every inverse frequency by 16. That stretches the usable position range but also changes short-distance rotations in every feature pair. YaRN instead blends unscaled and interpolated frequencies: fast-rotating pairs retain their original frequencies, slow pairs are stretched, and a transition band joins them.11
The bounds 18 and 35 are derived from the implementation's floor/ceil correction range with original length 8,192 and beta_fast=32, beta_slow=1. It locates feature-pair indices by how many rotations they complete over the original context: . At indices 0โ18 the global frequency stays unchanged; at 35โ63 it is divided by 16. The intervening pairs are blended.11
The checkpoint also specifies an attention factor . The implementation multiplies both sine and cosine by in global layers. Both rotated Q and K are consequently scaled, so the pre-softmax dot product acquires . Local layers use factor one. Reproducing only the frequency ramp while omitting this amplitude factor changes inference.4511
The report's rationale and the rotation identity support preserving local geometry. They do not prove that every long-context task succeeds or that this recipe dominates all alternatives. The original report's YaRN comparison also documents a QA prompt-formatting issue in its RULER evaluation. Those extension experiments concern Mellum2 training, not a fresh measurement of the released 2.1 checkpoint.10
Eight experts are computation, not an eight-expert model
For each normalized token vector , the router projects to 64 logits, computes softmax in float32, selects the eight largest probabilities and renormalizes the selected set. Each selected expert uses a SiLU-gated MLP with intermediate width 896. Its output is weighted and accumulated into the token's 2,304-dimensional result.5
Each expert has three matrices containing weights. Across 28 layers, all 64 experts contain about 11.10 billion weights, while the eight selected experts contribute about 1.39 billion weights to one token's path. Attention, routers, embeddings and the output projection are additional. These are derived matrix counts; โ2.5B activeโ is a rounded model descriptor, not a guaranteed dense-model speed equivalence.
The published tensor metadata reports 12,149,923,072 BF16 parameters overall, or about 22.63 GiB at two bytes each, before cache and runtime memory. Unselected experts remain stored unless a deployment explicitly offloads or partitions them. Different tokens in a batch can collectively activate many more than eight experts. vLLM routes this model through its fused MoE implementation; the reference Python expert loop describes semantics, not the GPU kernel's execution schedule.312
A concrete hybrid KV-cache budget
For one request, one layer and one retained position, BF16 K plus V require bytes. GQA already reduces this by eight relative to 32 separately cached KV heads. Sliding attention then limits the positions needed in 21 layers; the seven global layers continue to grow with sequence length.4
Using a 1,024-position local working window, the logical K/V payload is below. This assumes an unsharded model, BF16 cache, no prefix sharing and a runtime that retains only the required local history. It includes the current position and excludes allocator overhead.
| Sequence length | All 28 layers full, GQA | Hybrid local/global, GQA |
|---|---|---|
| 8,192 | 0.4375 GiB | 0.15039 GiB |
| 32,768 | 1.7500 GiB | 0.47852 GiB |
| 131,072 | 7.0000 GiB | 1.79102 GiB |
At the configured maximum, the hybrid payload is 74.41% smaller than the hypothetical all-full version with the same heads. This is a computed storage comparison, not a measured VRAM reduction or a comparison with a separately trained model. The seven global caches still hold 1.75 GiB; local layers add about 0.041 GiB. Multiplying by the number of independent requests estimates their logical payload, not total server VRAM.
The Transformers dynamic cache keeps the last 1,023 local positions for the next call and combines them with new tokens for attention. Its update returns the larger concatenated tensors, and its stored slices can share their backing allocation. In particular, a large prefill or chunk can have a larger physical footprint than the retained tensor shape suggests. Paged allocation, temporary buffers and static caches can add further space. A local attention mask alone does not establish eviction or actual memory usage.13
Even the idealized maximum-context budget is 22.63 + 1.79 โ 24.42 GiB for BF16 weights and KV alone. Activations, scratch space and serving reservations come on top. This explains why a โ2.5B activeโ model is not automatically comfortable in a 24 GiB budget. For allocation and concurrency details, see KV-cache memory math and quantization.
Serving the released path
The pinned vLLM implementation selects rope_parameters[layer_type] and passes the per-layer window into attention. It also handles the explicit 128-wide heads. A runtime that merely loads expert weights but applies one global RoPE configuration to all layers does not reproduce this checkpoint.6
With vLLM 0.31.0 already installed, the following starts the pinned BF16 model at a deliberately smaller 32,768-token limit. The parser flags follow the model card; they separate reasoning output and parse tool calls, rather than executing tools. Raise the context limit only after validating the memory budget and workload.2
vllm serve JetBrains/Mellum2.1-12B-A2.5B-Thinking \
--revision 92ddae9fc7665e9f801d141d2e5a6b2caf2460c4 \
--dtype bfloat16 \
--max-model-len 32768 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser hermesThis command is a source-checked configuration example, not an independently executed GPU test. Verify the selected attention backend supports the mixed window pattern and examine actual cache reservations. Tensor parallelism above four ranks can replicate KV heads in the vLLM implementation, so aggregate memory does not follow a naive division by rank count.6
GGUF lowers weight storage separately from the attention design. The official Q4_K_M file is listed as 8.1 GB; that is a file size, not total working memory. Its reported quantization comparisons use Wikitext-2 with a 512-token context. Those logit comparisons do not establish agentic coding accuracy or quality at 128K, and quantized weights do not automatically imply a quantized KV cache.7
Mellum2.1 provides a concrete example of composing sparse feed-forward computation with cheaper local attention while retaining periodic global access. Layer-selective YaRN adapts the global path's positional geometry. Whether that balance suits a coding service depends on its tasks, request lengths, quantization and runtime; architecture and parameter counts alone cannot establish throughput or answer quality.
Primary sources
Footnotes
-
JetBrains, Mellum2.1 Gets to Work, October 2026; checked 2026-10-11. Explicitly distinguishes post-training changes from unchanged architecture. โฉ
-
JetBrains, Mellum2.1 Thinking model card, pinned revision. โฉ โฉ2 โฉ3
-
Hugging Face, JetBrains checkpoint metadata at the inspected revision, BF16 tensor count; checked 2026-10-11. โฉ โฉ2
-
JetBrains, released configuration. โฉ โฉ2 โฉ3 โฉ4 โฉ5 โฉ6
-
Hugging Face Transformers 5.10.2, Mellum implementation:
MellumAttention,MellumRotaryEmbedding,MellumTopKRouter,MellumExperts,MellumDecoderLayer. โฉ โฉ2 โฉ3 โฉ4 โฉ5 โฉ6 -
vLLM 0.31.0, Mellum implementation, per-layer rotary/window configuration and tensor-parallel KV-head allocation. โฉ โฉ2 โฉ3
-
JetBrains, official GGUF model card and files, pinned revision; card lists sizes and quantization measurement conditions. โฉ โฉ2
-
JetBrains, released weight index. โฉ
-
Transformers 5.10.2, mask implementation,
sliding_window_overlay,sliding_window_causal_mask_function. โฉ -
Kojic et al., Mellum 2 Technical Report, 2026-05-29, especially Sections 2.2, 4.1 and C.1. Describes the inherited architecture and limits of its original extension evaluation. โฉ โฉ2
-
Transformers 5.10.2, YaRN implementation,
_compute_yarn_parameters. โฉ โฉ2 โฉ3 -
vLLM 0.31.0, Qwen3 MoE implementation reused by Mellum,
Qwen3MoeSparseMoeBlockandFusedMoEFactory. โฉ -
Transformers 5.10.2, cache implementation,
DynamicSlidingWindowLayer.updateandDynamicCache. โฉ