Research

Kolibri-1 Released Under Apache 2.0: 78B MoE Runs on a Single GPU and Outperforms Peer Models in German

Heidelberg-based AI lab Aleph Alpha published Kolibri-1, an open-weight mixture-of-experts model spanning 78 billion total parameters with only 3.46 billion active per token. The architecture runs on a single enterprise GPU, delivers a 1-million-token context window, and beats same-size open alternatives across German and EU regulatory evaluations.

By Julian Thorne · 2026-10-06 · 11 min read

Aleph Alpha open-sourced Kolibri-1 on October 5, 2026, releasing model weights under the permissive Apache 2.0 license. The release gives European enterprises and international researchers an open foundation model designed specifically for data sovereignty and regional legal compliance.

Kolibri-1 uses a 78-billion-parameter Mixture-of-Experts (MoE) architecture. Instead of routing tokens through all 78 billion weights, the gating network activates only 3.46 billion parameters per forward pass across its top-2 expert paths.

This sparse compute design changes deployment economics: full-precision FP8 inference runs smoothly on a single 80 GB NVIDIA H100 or A100 GPU, eliminating the networking latency and multi-GPU interconnect overhead that slows down standard 70B dense models.


Hardware Footprint: Single-GPU Execution Profiles

Running standard dense models like Llama 3.3 70B requires either an 8-GPU node for low-latency batching or aggressive 4-bit quantizations that degrade factual recall. Because Kolibri-1 calculates activations through only 3.46 billion parameters at any step, its compute requirements remain light.

┌────────────────────────────────────────────────────────────────────────┐
│                      Kolibri-1 Single-GPU Memory Map                   │
├──────────────────────────┬──────────────────────────┬──────────────────┤
│ Quantization Format      │ Target Hardware          │ Memory Bandwidth │
├──────────────────────────┼──────────────────────────┼──────────────────┤
│ FP8 (Native vLLM)        │ 1x NVIDIA H100 80GB SXM  │ 84 tokens/sec    │
│ INT8 SmoothQuant         │ 1x NVIDIA A100 80GB PCIe │ 56 tokens/sec    │
│ AWQ 4-Bit                │ 1x NVIDIA RTX 4090 24GB  │ 42 tokens/sec    │
│ Q4_K_M (llama.cpp)       │ Apple M3 Max (36GB RAM)  │ 31 tokens/sec    │
└──────────────────────────┴──────────────────────────┴──────────────────┘

The small active footprint allows engineering teams to host dedicated sovereign instances on internal corporate servers, avoiding third-party American cloud APIs.


German & European Benchmark Results

The core technical differentiator is native linguistic representation. Standard frontier foundation models allocate less than 3% of their training tokens to German text, forcing tokenizers to break common compound words like Bundesdatenschutzgesetz into ten or more disjointed fragments.

Aleph Alpha designed a bilingual German/English vocabulary of 131,072 tokens and allocated 21.3% of all pre-training compute to vetted German-language publications, technical documentation, and statutory texts.

Benchmark Dataset Evaluation Task Llama 3.3 70B (Dense) Mistral Large 2 (123B) Kolibri-1 (3.46B Active)
German MMLU (de-MMLU) High-school & College Exams 74.2% 76.8% 79.4%
German Legal QA (GerLaw) Statutory code interpretation 68.1% 71.5% 83.2%
GSM8K Math (Reasoning Mode) Multi-step arithmetic 86.4% 89.1% 88.7%
ToolBench Action Success API function call accuracy 78.5% 82.0% 84.3%
Context Retention (1M tokens) Needle-in-a-Haystack pass 92.0% 88.4% 97.8%

On German statutory analysis (GerLaw), Kolibri-1 outperforms Mistral Large 2 by 11.7 percentage points, despite executing less than 5% of the active parameter calculations per token.


Regulatory Architecture: EU AI Act and GDPR

Aleph Alpha trained Kolibri-1 with explicit European Union legal compliance built into the data pipeline:

  1. Copyright Filtering: All pre-training documents were matched against EU digital copyright registries, excluding scrapers that ignored robots.txt machine-readable reservations under Article 4(3) of the EU DSM Directive.
  2. GDPR Privacy Cleansing: The training pipeline removed personally identifiable information, phone numbers, and addresses using regex-guided named entity recognition models.
  3. General-Purpose AI Code of Practice: Model cards and weights release full training compute logs and carbon emission figures, satisfying upcoming EU regulatory thresholds for foundation models.