Aleph Alpha Drops Kolibri: 78B Mixture-of-Experts Open Model with 1M Context Window and Apache 2.0 License
Heidelberg-based AI lab Aleph Alpha published open weights for Kolibri, a 78-billion-parameter Mixture-of-Experts model activating 3.46B parameters per token, featuring a 1-million-token context length and native EU AI Act compliance.
German artificial intelligence research laboratory Aleph Alpha published open weights and source code for Kolibri, a 78-billion-parameter Mixture-of-Experts (MoE) model.
Licensed under Apache 2.0, Kolibri activates just 3.46 billion parameters per token, delivering inference speeds competitive with lightweight 3B models while maintaining the parametric knowledge and reasoning breadth of large models.
Kolibri stands alongside Mistral as one of the few high-capacity, general-purpose open-weight foundation models produced in the European Union.
Architectural Specifications: Kolibri-78B-MoE
The table below contrasts Kolibri against leading open-weight Mixture-of-Experts architectures:
| Specification Metric | Aleph Alpha Kolibri | Mistral Mixtral 8x7B | Qwen-2.5-57B-A14B |
|---|---|---|---|
| Total Parameters | 78.2 Billion | 46.7 Billion | 57.1 Billion |
| Active Parameters / Token | 3.46 Billion | 12.9 Billion | 14.2 Billion |
| Expert Topology | 32 total experts (Top-2 routing) | 8 total experts (Top-2 routing) | 64 total experts (Top-8 routing) |
| Context Window | 1,048,576 tokens (1M) | 32,768 tokens (32k) | 131,072 tokens (128k) |
| Pre-Training German Data | 21.3% | ~3.8% | <1.0% |
| License | Apache 2.0 | Apache 2.0 | Apache 2.0 |
| Inference Footprint (FP8) | ~42 GB VRAM (Fits on 1x A100) | ~48 GB VRAM | ~60 GB VRAM |
Sparse Activation & Long-Context Routing
Kolibri achieves high compute efficiency by decomposing its feed-forward layers into 32 granular routed experts plus 2 shared experts. The routing gate selects the top 2 routed experts per token, keeping the memory bandwidth requirement to 3.46B parameters:
[Input Tokens: Up to 1,000,000 Length]
│
▼
┌────────────────────────────────────────────────────────┐
│ FlashAttention-3 & RoPE Positional Encoding │
└──────────────────┬─────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────┐
│ Mixture-of-Experts Layer │
│ Shared Experts (2) │ Top-2 Gated Routed Experts │
│ (Universal syntax) │ (Domain/Language specialists) │
└──────────────────┬─────────────────────────────────────┘
│
▼
[3.46B Active Params / Forward Pass]
The attention mechanism utilizes grouped-query attention (GQA) with an 8:1 ratio and dynamic rotary position embeddings (RoPE) scaled to handle sequence lengths exceeding 1 million tokens without attention degradation.
Built for the EU AI Act and GDPR
Commercial adopters in Germany, France, and Switzerland face strict compliance obligations under the EU Artificial Intelligence Act and the General-Purpose AI (GPAI) Code of Practice.
Aleph Alpha designed Kolibri with verifiable provenance:
- Audited Training Provenance: The dataset excludes non-compliant web scrapes, adhering to EU copyright exceptions and opt-out directives.
- Organic European Language Tokenizer: Pre-training incorporated 21.3% native German source texts (legal statutes, technical engineering manuals, clinical records), avoiding compression penalties inherent in English-biased tokenizers.
- Explicit Reasoning Tokens: Includes native special tokens (
<reasoning_start>,<reasoning_end>) allowing enterprise safety arbiters to inspect internal chain-of-thought traces before final answers stream to users.
How to Run Kolibri via vLLM and Transformers
Weights are available on Hugging Face under Aleph-Alpha/Kolibri-1:
# Install latest vLLM with Kolibri MoE kernel support
pip install --upgrade vllm
# Launch high-throughput OpenAI-compatible server
python -m vllm.entrypoints.openai.api_server --model Aleph-Alpha/Kolibri-1 --tensor-parallel-size 2 --max-model-len 131072 --quantization fp8
Prompting with explicit reasoning:
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("Aleph-Alpha/Kolibri-1")
model = AutoModelForCausalLM.from_pretrained("Aleph-Alpha/Kolibri-1", device_map="auto")
inputs = tokenizer("<|im_start|>system
Mode: Explicit Reasoning<|im_end|>
<|im_start|>user
Analyze German GDPR Article 28 data processor obligations.<|im_end|>
<|im_start|>assistant
", return_tensors="pt").to("cuda")
output = model.generate(**inputs, max_new_tokens=1024)
print(tokenizer.decode(output[0]))
Summary
Kolibri establishes that European artificial intelligence labs can produce world-class open-weight architectures that respect rigorous regional regulatory statutes while outperforming competing models in parameter activation efficiency.