Kolibri 1 is a released English-German open-weight language model, not an architecture preview. Aleph Alpha published FP8 and BF16 checkpoints under Apache 2.0 on 3 October 2026, together with a model card, a 189-page technical report and an inference-only vLLM plugin. The released model has 78.1 billion total parameters and activates a reported 3.46 billion per token. This article describes those released artifacts; benchmark and throughput results remain vendor measurements.123
The interesting part is not simply that Kolibri is a Mixture of Experts. Every one of its 50 blocks combines Grouped-Query Attention (GQA) with an MoE. Forty blocks use a short sliding window and RoPE, while ten periodically inserted global blocks use full causal attention without positional encoding. Every MoE evaluates one shared expert and six of 384 routed experts. The result trades most long-range attention work for local attention without eliminating either the growing global KV cache or the resident 78B weight set.
The released stack
| Component | Checkpoint value | Runtime consequence |
|---|---|---|
| Residual stream | 50 blocks, width 2,560 | Each block consumes and returns [B, S, 2560]. |
| Attention | 48 query heads, four KV heads, head dimension 128 | The logical Q/K/V widths are 6,144/512/512; GQA shares each KV head across 12 query heads. |
| Hybrid pattern | 40 sliding-window blocks, ten full-attention blocks | Four local blocks precede each global block. The local window retains 512 preceding tokens plus the current token. |
| Positional path | RoPE with base 10,000 only in local blocks; no position encoding in global blocks | Global attention avoids a learned or rotated absolute-distance limit, but it still incurs full-context cache and compute. |
| MoE per block | 384 routed experts, top-6, plus one shared expert | Seven expert MLPs execute for each token; all 385 sets of expert weights must remain available. |
| Expert shape | SwiGLU, 2560 -> 512 -> 2560 | Experts are deliberately narrow; each has two input projections and one output projection. |
| Released precision | FP8 E4M3 block-quantized checkpoint and BF16 checkpoint | The compact checkpoint is a mixed-precision runtime artifact, not eight-bit storage for every tensor. |
The block uses sandwich normalization: RMSNorm before and after both attention and MoE, with residual additions after each sublayer. Queries and keys also receive per-head RMSNorm. In logical, unsharded form, attention therefore maps [N, 2560] to Q, K and V tensors of [N, 48, 128], [N, 4, 128] and [N, 4, 128]. The 6,144-wide attention output is projected back to the 2,560-wide residual stream. These are checkpoint shapes; tensor parallelism partitions them in the serving implementation.45
Local RoPE plus global NoPE attention
In the 40 local blocks, RoPE rotates normalized queries and keys and the causal mask exposes at most 512 earlier tokens. Decode work and retained KV state in these blocks stop growing once the window is full. In every fifth block, the window disappears and the token may attend to the complete prefix. Those ten blocks apply no RoPE at all, a pattern often called RNoPE for the global path. Information can therefore cross windows through the global layers even though most layers remain local.25
This is not linear attention. For a prefill of length S, the local layers perform attention over roughly S × 513 positions, but the ten global layers still have quadratic S² attention matrices. During decoding, every new token reads the complete retained cache in those global layers. The design reduces the coefficient of global attention by four fifths; it does not remove the long-context bottleneck.
The logical cache floor makes the distinction concrete. One token contributes K and V for four heads of width 128. Across ten full-attention layers that is
cached values per retained token: 20 KiB in BF16 or 10 KiB in FP8. At 262,144 tokens the global portion alone is 5 GiB in BF16 or 2.5 GiB in FP8 per request; at 1,048,576 tokens it is 20 GiB or 10 GiB. The 40 full local windows add about 40.1 MiB in BF16 or 20.0 MiB in FP8. These are arithmetic lower bounds derived from published tensor dimensions, before block tables, alignment, allocator fragmentation and workspace. The KV-cache guide explains why an engine's real reservation is larger.
Routing: selection bias and mixture weights are separate
For a flattened token matrix x of shape [N, 2560], the router produces FP32 logits [N, 384]. The released plugin implements a subtle two-step rule:
- Add a learned expert-correction bias to the logits and select the top six expert IDs.
- Gather the unbiased logits for those IDs, apply sigmoid, and use the six scores as mixture weights without renormalizing them to sum to one.
The correction bias can therefore alter which experts receive a token without directly scaling the selected expert outputs. The technical report links those stored biases to Exact Quantile Balancing during training. At inference, routing remains token-choice top-k: it does not impose an exact per-batch capacity on each expert. Real expert-parallel deployments must still handle skew and the all-to-all traffic needed to send tokens to the devices that own their selected experts.25
Each routed or shared expert is a SwiGLU MLP. Two 2560 -> 512 projections form SiLU(gate) × up, followed by a 512 -> 2560 projection. Ignoring biases, one expert therefore contains
matrix weights. Six routed experts plus the always-active shared expert touch about 27.5 million expert-matrix weights per block and token. All 385 experts contain about 1.51 billion such weights per block, however. Sparse activation lowers arithmetic and weight traffic for a token, but it does not turn the checkpoint into a 3.46B storage problem. Weight residency, expert placement and interconnect bandwidth remain independent capacity constraints.
The FP8 path is explicitly mixed precision
The main checkpoint stores eligible linear weights in FP8 E4M3 using 128 × 128 blocks with an FP32 scale per block. Activations are quantized dynamically in groups of 1 × 128. Embeddings, normalization layers, the LM head and MoE router remain BF16; the serving path keeps router output logits in FP32. The official launch command also selects an FP8 KV cache. Consequently, “FP8 model” does not mean every value occupies one byte, and the reported roughly 78 GB weight footprint does not include per-request KV cache, engine workspace or duplicated/sharded runtime state.12
There is also a concrete kernel constraint. The official plugin rejects vLLM's DeepGEMM E8M0 scale path for this checkpoint: Kolibri carries FP32 block scales, while that path would round them to powers of two. The implementation instructs operators to disable VLLM_USE_DEEP_GEMM_E8M0. This is a compatibility requirement from the released code, not a general statement that DeepGEMM or FP8 is inaccurate.5
Aleph Alpha publishes a container and the aleph-alpha-inference package, currently tied to vLLM 0.29. The documented command serves Aleph-Alpha/Kolibri-1 with --kv-cache-dtype fp8; the BF16 checkpoint uses Aleph-Alpha/Kolibri-1-BF16 and omits that flag. The model card lists about 78 GB for FP8 weights and a minimum of one H200/B200/B300 or two 80 GB A100/H100 GPUs. Those are vendor deployment minima, not guarantees for a chosen context length or concurrency. The VRAM guide covers the additional memory terms.13
Native 256K and evaluated 1M are different claims
Kolibri pre-trained at 16,384 tokens, mid-trained at 65,536, and received a final long-context stage at 262,144 tokens. Its configuration correspondingly defaults to max_position_embeddings: 262144. Because only the bounded local layers use RoPE and global layers use no positional encoding, the runtime can extend the sequence without changing a RoPE scale. The official recipe reaches 1,048,576 tokens by overriding the model length and configuration.143
That makes one million tokens an evaluated extrapolation, not the native training length. The report evaluates RULER through 1M and explicitly finds task-dependent degradation: some subtasks remain stable while others decline. Aleph Alpha recommends at most 262,144 tokens for efficiency and complex tasks. Raising the limit also grows the ten global KV caches to the sizes calculated above and leaves quadratic global prefill work. “No positional rescaling required” is therefore not the same as “no quality, latency or memory cost.”
Kolibri 1 is production-usable in the practical sense relevant here: weights, license, configuration, BF16 fallback, container and a pinned serving implementation are available. Its operational trade is equally concrete. Hybrid attention caps most caches but preserves ten expensive global layers; narrow sparse experts reduce active compute but require a 78B resident pool and routing communication; FP8 cuts storage and cache bytes but depends on a specific mixed-precision kernel path. Capacity planning must evaluate all three together rather than treating the 3.46B active-parameter number as a hardware requirement.
Sources
Footnotes
-
Aleph Alpha, official
Kolibri-1model card and released FP8 weights. ↩ ↩2 ↩3 ↩4 -
Aleph Alpha, Kolibri: A Sovereign European Model on the Pareto Frontier, technical report, October 2026. ↩ ↩2 ↩3 ↩4
-
Aleph Alpha, official
aleph-alpha-inferencevLLM plugin and serving instructions. ↩ ↩2 ↩3 -
Aleph Alpha, pinned Kolibri 1 vLLM implementation,
kolibri1.py. ↩ ↩2 ↩3 ↩4