Aleph Alpha Releases Kolibri, an Open-Weight 78.1B MoE Model
The German AI lab releases an open-weight mixture-of-experts model optimized for bilingual enterprise workloads, long contexts, and sovereign European compliance.

Key takeaways
- Kolibri features 78.1 billion total parameters but activates only 3.46 billion parameters per token across 50 transformer layers.
- The model supports context lengths up to 1,048,576 tokens, with native training at 262,144 tokens using hybrid sliding-window and full attention.
- Released under an Apache 2.0 license, Kolibri is engineered specifically for English and German enterprise and public sector applications.
- Training utilized 24 trillion tokens across clusters in Germany and Finland with compliance tailored to the EU AI Act and GDPR.
On October 3, 2026, German AI laboratory Aleph Alpha released Kolibri, an open-weight mixture-of-experts (MoE) reasoning model developed for German and English. Kolibri carries 78.1 billion total parameters but activates only 3.46 billion per token during inference, balancing dense model capacity with lower operational compute costs. Published under the permissive Apache 2.0 license, the model provides an open-weight alternative tailored for regulated European sectors that demand sovereign infrastructure and strict regulatory alignment.
According to Aleph Alpha's model card on Hugging Face, Kolibri is designed to perform structured data extraction, multi-step reasoning, coding, and retrieval-augmented generation (RAG) across long contexts. The architecture supports up to 1,048,576 tokens (1 million tokens), though the developers recommend serving sequences at or below 262,144 tokens for latency-sensitive or complex workloads.
Architecture and Tokenizer Design
Kolibri's underlying architecture comprises 50 transformer layers with a model width of 2,560. According to technical specifications detailed by MarkTechPost, each MoE layer features 384 routed experts alongside one shared expert. Tokens are evaluated via a sigmoid router and dispatched to the top six routed experts plus the single shared expert, activating 4.4% of total parameters per token. Load balancing during pre-training was handled using Exact Quantile Balancing and Load-Error Injection.
To keep compute and memory manageable over extended context windows, Kolibri uses a 4:1 hybrid attention mechanism. Forty of the 50 layers employ sliding-window attention constrained to the preceding 512 tokens using Rotary Position Embeddings (RoPE), maintaining a fixed-size Key-Value (KV) cache. The remaining 10 layers (every fifth layer) execute full grouped-query attention (GQA) across the entire sequence length without positional encodings. Aleph Alpha notes that this hybrid attention scheme allows the model to process sequences four times longer than a standard full-attention architecture at matching compute budgets.
The model includes a custom UniBPE tokenizer with a vocabulary of 128,000 tokens. As reported by MarkTechPost, the tokenizer achieves an efficiency of 4.90 bytes per token on German text, compared to 4.35 bytes per token for the GPT-5 tokenizer, yielding 11.2% fewer tokens on German web data while maintaining 4.58 bytes per token on English.

Iterative Pipeline and the Kolibri Origin Foundation
Kolibri’s development was preceded by Kolibri Origin, an internal 30.6-billion-parameter research model with 3.27 billion active parameters and a 65,536-token context window that was completed on June 11, 2026, but never released publicly. As explained on Aleph Alpha's technical blog, Origin served as an end-to-end testbed to stabilize automated continuous integration, data shuffling, and ablation workflows.
Over the three months between the completion of Origin and Kolibri's final pre-training on September 11, 2026, the lab scaled total parameters from 30.6B to 78.1B, expanded expert counts per layer from 128 to 384, and broadened context limits to 1M tokens. Total training volume increased from 7.51 trillion to 24 trillion tokens, comprising 20T pre-training tokens, 3.44T mid-training tokens at a 65,536-sequence length, and 201B tokens dedicated to extending context to 262,144 tokens.
Pre-training was conducted on 768 NVIDIA B200 GPUs across 96 HGX nodes over 21 days (392,000 GPU-hours), generating an estimated total training energy consumption of 950 MWh, according to Aleph Alpha's Hugging Face repository. The training corpus contained approximately 62.5% English, 23.9% German (over 2 trillion organic German tokens), and 13.6% source code, maintaining a knowledge cutoff date of June 18, 2026.
Post-training combined supervised fine-tuning (SFT) using MergeMix with reinforcement learning across 1.2 million internal environments. Aleph Alpha also integrated its Merlin-Arthur protocol, training the model to intentionally abstain and output "I don't know" when retrieved context lacks sufficient evidence.
Benchmark Performance and Comparisons
Benchmark evaluations reported by MarkTechPost place Kolibri at the top of its active-parameter class across several core benchmarks:
- Mathematics: Kolibri scored 96.9 on AIME 2025 (87.5 in German) and 96.0 on AIME 2026 (90.0 in German), outpacing rivals like Qwen3.6 35B-A3B (84.6 on AIME 2025) and NVIDIA Nemotron 3 Super (91.7 on AIME 2025).
- Knowledge and Reasoning: On English GPQA Diamond, Kolibri achieved 84.3 (81.3 in German), leading Qwen3.6 (83.4) and Nemotron 3 Super (78.0).
- Agentic Workflows: Kolibri averaged 63.4 on English agentic tests and scored 61.4 on the Berkeley Function Calling Leaderboard (BFCL) v4, trailing Qwen3.6's 67.2.
- Overall Standings: Kolibri recorded overall aggregate benchmark scores of 75.5 in English and 70.8 in German among 12 evaluated MoE models.
According to an analysis by hyper.ai, Kolibri exhibits strong mathematical reasoning and enterprise RAG grounding, but shows comparative weaknesses in closed-book general knowledge retrieval, multi-turn tool calling, and mid-range context accuracy relative to dense alternatives.

Hardware Specs and Local Deployment
Kolibri's FP8 weights (formatted as float8_e4m3fn in 128×128 blocks with bfloat16 embeddings and router layers) present a storage footprint of roughly 78 GB. MindStudio's technical overview highlights that operational deployment requires additional VRAM to accommodate activations and the FP8 KV cache, with real-world tests measuring around 140 GB of VRAM utilization on dual-GPU configurations.
Aleph Alpha lists minimum serving hardware as two NVIDIA A100 80GB, two H100 SXM5, or a single H200, B200, or B300 GPU. Recommended hardware includes two H100 SXM5, two H200, or a single B200 or B300 GPU. The model is served via vLLM using the dedicated aleph-alpha-inference package:
pip install 'aleph-alpha-inference>=1'
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choice
Kolibri incorporates configurable reasoning modes. Developers can adjust reasoning intensity by supplying none, low, medium, or high via the chat template, or disable thinking entirely to reduce latency.
European Sovereignty and Enterprise Alignment
Aleph Alpha emphasizes European data sovereignty and regulatory compliance as foundational design priorities for Kolibri. As reported on Aleph Alpha's corporate blog, the model was constructed end-to-end within Germany and Finland on European compute infrastructure, ensuring full traceability from data ingestion through pre-training.
As a signatory to the European Union's General-Purpose AI (GPAI) Code of Practice, Aleph Alpha designed Kolibri to align with the EU AI Act and GDPR guidelines. Pre-training data pipelines automatically redacted personal data, and the model's transparent Apache 2.0 open-weight release provides public administration, aerospace, manufacturing, and financial institutions with the ability to self-host without routing proprietary text through external third-party APIs.
Frequently asked questions
What is Kolibri's parameter size and architecture?
Kolibri is a mixture-of-experts model containing 78.1 billion total parameters. It dynamically routes tokens across 384 routed experts and one shared expert per layer, activating approximately 3.46 billion parameters per token.
What context length does Kolibri support?
Kolibri natively supports 262,144 tokens and has been validated up to 1,048,576 tokens (1M context window) using a hybrid sliding-window and full grouped-query attention structure.
Under what license is Kolibri released?
Kolibri is released under the open-source Apache 2.0 license, allowing free commercial use, modification, and on-premise deployment.
What hardware is required to run Kolibri?
Serving the FP8 model checkpoint requires a minimum of two NVIDIA A100 80GB, two H100 SXM5, or a single H200/B200/B300 GPU, with dual H100 SXM5 or H200 units recommended to accommodate KV caching overhead.
Sources
- Aleph-Alpha/Kolibri-1Hugging Face · Oct 2, 2026 · Official
- Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parametersmarktechpost.com · Oct 4, 2026
- Kolibri Has Landed: A Sovereign Open-Weight Model — Aleph AlphaAleph Alpha · Oct 3, 2026
- Kolibri-1: Aleph Alpha's Open-Weight German-English MoE ModelMindStudio · Oct 4, 2026
- Aleph Alpha Releases Kolibri 1: Sovereign German MoE LLMhyper.ai
How this story was made: the newsroom picked it up from Google News, gathered the full text of the sources above, and drafted it with AI assistance. Every factual claim was then checked against those sources before publishing (33 claims checked). Illustrations marked as AI-generated are not photographs. Spotted an error? Tell us.
Published October 5, 2026 at 01:15 UTC


