Aleph Alpha releases Kolibri
Sovereign AI made in Germany
Heidelberg, 5 October 2026 - Aleph Alpha has released Kolibri, its specialized large language model for mission-critical applications in public administration and industry. Developed and trained in Europe, Kolibri supports German and English. The model has been available since 3 October, and is designed for deployment on infrastructure customers control, with documented data provenance and design decisions.
Key features of Kolibri:
- Efficient architecture: Kolibri combines 78 billion total parameters with around 3 billion active per token. Its mixture-of-experts architecture selectively activates parts of the model during inference.
- Built for German: Kolibri is trained to reason in German, with German-language text accounting for around 23% of its pre-training data.
- German-optimized tokenizer: Kolibri’s tokenizer is tailored to German and designed to represent German text with fewer tokens, supporting efficient processing of German-language content.
- Documented compliance measures: The technical report details the measures taken to address copyright, data protection and EU AI Act requirements, including training-data screening and checks on third-party datasets.
Ilhan Scheer, CEO of Aleph Alpha: "Kolibri demonstrates that we have the talent and expertise in Germany to develop competitive AI models. For us, AI sovereignty means freedom of choice by retaining the ability to build and advance this technology, and giving customers control over how they use it. Our ambition is to bring that capability into everyday operations in industry and public administration."
Samuel Weinbach, Co-founder and Co-Chief Research Officer: "We designed Kolibri around the needs of German-language applications, from a tokenizer optimized for German to training the model to reason in German. For our customers, this is an important step towards making AI useful in everyday workflows."
Kolibri supports agentic workflows, retrieval-augmented generation (RAG) and native tool calling. In RAG applications, the model uses supplied documents to answer questions and is trained to abstain when those documents do not contain sufficient evidence. Built-in reasoning controls allow users to adjust the balance between reasoning effort and response speed.
Aleph Alpha’s approach to training-data governance is documented in Kolibri’s technical report, including measures addressing copyright and data protection. All training data was screened against a blocklist of more than 4.5 million URLs, drawing on sources including the European Commission’s Piracy Watch List. Each third-party dataset was assessed for license terms, lawful sourcing and adherence to opt-outs before use.
Kolibri was developed by Aleph Alpha. The company recently announced an agreement with Cohere to build a transatlantic sovereign AI company. The transaction remains subject to regulatory approval. Until closing, Aleph Alpha continues to operate independently.
The model, its applicable license terms and supporting documentation are available on Hugging Face. Further information on Kolibri’s architecture, evaluations and training methodology is available in the Tech Report and our blog post.
Contact: press@aleph-alpha.com
About Aleph Alpha
Aleph Alpha was founded in 2019 with the mission to research and build sovereign, human-centric AI for a better world. With an international team of scientists and engineers, the company researches and develops Specialized Large Language Models. Its co-created solutions are the first choice for companies and government institutions that want to maintain sovereignty, secure data, and develop trustworthy applications. With its headquarters in Heidelberg, Aleph Alpha employs some 200 talent across four locations in Germany. For more information, visit aleph-alpha.com.