Kolibri-1 logo

Kolibri-1

Visit

Aleph Alpha’s open-weight reasoning model for German and English, with 78B total parameters, tool calling and long-context inference.

Share:
View alternatives

Kolibri-1

Kolibri-1 is Aleph Alpha’s open-weight text reasoning model, released on October 3, 2026. It targets German and English work such as document analysis and tool-assisted tasks. Its mixture-of-experts architecture has about 78 billion total parameters and activates about 3.46 billion per token. That distinction matters: active parameters describe computation, while deployment still requires storing the full weights.

Features and suitable work

The model supports explicit reasoning and tool calling. Its advertised maximum context is 1,048,576 tokens, but the official model card recommends staying at or below 262,144 tokens for serving efficiency and complex tasks. The native long-context training length is 262,144 tokens. Use the larger limit as something to test rather than a guarantee that an entire document will be handled reliably.

A German-English internal knowledge assistant is a useful pilot: retrieve a small set of approved passages, ask for a sourced answer, and verify every cited passage. Another practical application is a narrowly scoped tool workflow with allowlisted functions. Tool-call syntax support does not establish that an action is safe or correct.

Hardware, licensing and cost

The main repository distributes FP8 weights under Apache 2.0. The card estimates approximately 78 GB for model weights and lists minimum configurations including two 80 GB A100 GPUs, two H100 SXM5 GPUs, or one H200, B200 or B300. Context caches and serving overhead need additional capacity. A separate BF16 checkpoint has different requirements; do not apply the FP8 estimates to it.

There is no model-weight subscription fee under this license, but GPU hosting, storage, operations and evaluation still cost money. This entry does not quote an unverified hosted API price.

Getting started

  1. Read the model card’s serving instructions and confirm that your runtime supports this checkpoint and precision.
  2. Start with a short German-English test set and a conservative context cap.
  3. Compare direct answers with retrieval-assisted answers on the same questions.
  4. Add tool calls only after validating parsing, permissions and failure handling.
  5. Measure memory, latency and answer accuracy under your intended concurrency before scaling.

FAQ and comparisons

Is this a lightweight 3B model? No. The active parameter count is approximately 3.46B, but the total is approximately 78B and the full checkpoint must be resident in the deployment’s memory arrangement.

Does the million-token limit remove the need for retrieval? No. Retrieval can reduce cost and focus evidence, and long-context quality should be tested separately from accepted input length.

Compare Qwen3.8-Flash-Next when selecting another MoE model. Ollama is a runner rather than a substitute checkpoint; check compatibility before assuming it can serve Kolibri. Browse other models and open-weight tools.

Start with the official card and a small reproducible evaluation. Kolibri-1 is most relevant when German-English reasoning and control over deployment justify server-scale resources.

Comments

No comments yet. Be the first to comment!