
Aleph Alpha releases Kolibri with open weights and support for 1M-token contexts
The AMW Read
Kolibri meaningfully updates Aleph Alpha's foundation-model offering through a deliberate Apache 2.0 open-weight release and bilingual specialization, although segment-level impact depends on independent performance and deployment validation.
Aleph Alpha releases Kolibri with open weights and support for 1M-token contexts
Aleph Alpha has released Kolibri, an English-German model targeting government and regulated industries, with downloadable weights under Apache 2.0. The Mixture-of-Experts model has 78.1 billion total parameters and activates 3.46 billion per token. It supports tool calling, four reasoning settings, and contexts up to 1,048,576 tokens. Customers can deploy it on-premises using Aleph Alpha's inference package and a Kolibri-specific vLLM plugin. The company built the model in Germany and trained it in Germany and Finland.
The release positions Aleph Alpha in the foundation-model market around deployment control and bilingual specialization. Open weights let customers operate the model on their own hardware, while German represents 21.3% of its pre-training tokens. This gives buyers evaluating sovereign AI a concrete alternative to sending internal data to a third-party inference service. Its architecture combines sparse expert activation with sliding-window attention in 40 of 50 layers to contain serving costs. Aleph Alpha claims competitive quality relative to serving cost, but the reported math, coding, and tool-use results are vendor-run benchmarks, with the highest reasoning setting used where applicable.
For builders, the immediate task is to validate long-context reliability and deployment economics on their own workloads. Long-context adaptation reached 256,000 tokens; the one-million-token ceiling comes from the supplied serving settings, so maximum supported length should not be treated as evidence of equally reliable retrieval across that entire window. Teams should also test abstention and document grounding before relying on outputs in mission-critical workflows. The 3.46 billion active parameters describe computation per token, while the full model still has 78.1 billion parameters to accommodate in deployment planning.