Kolibri AI: What It Is, How It Benchmarks, And How It Compares With The Local Models On GoldieBench
By Julian Goldie · 2026-10-07 · GoldieBench Blog
Kolibri AI is a free, open-weight model from the German company Aleph Alpha, released on 3 October 2026 with 78.1 billion parameters and about 3.46 billion active per token.
People also search for it as Colibri, but the official name is Kolibri, which is German for hummingbird.
The first question everyone asks me is simple: is it any good compared with the models we already know?
I run GoldieBench, where 40 models have now produced 1,609 one-shot demos across 50 tasks.
Kolibri is not on the board yet, and I'll say that plainly up front.
So in this post I'll show you what Kolibri is, what Aleph Alpha's own benchmarks say, what it really takes to run, and how it lines up against the local and open models that are on the board.
What Kolibri AI is
Kolibri AI is Kolibri 1, Aleph Alpha's mixture-of-experts reasoning model with a focus on German and English.
Aleph Alpha is an AI company based in Heidelberg, and the model lives on Hugging Face as Aleph-Alpha/Kolibri-1.
Here is the spec sheet from the official model card.
| Spec | Official detail |
|---|---|
| Released | 3 October 2026 |
| Licence | Apache 2.0, and the repo is not gated |
| Total parameters | 78.1 billion |
| Active per token | About 3.46 billion |
| Experts | 384 per layer, with 6 routed and 1 shared, across 50 layers |
| Languages | German and English, both native |
| Context | 262,144 tokens native, validated up to 1,048,576 |
| Training | About 20 trillion pre-training tokens plus 3.44 trillion mid-training and 201 billion long-context tokens |
| Knowledge cutoff | 18 June 2026 |
| Features | Reasoning mode with effort levels, plus tool calling |
| Input and output | Text only |
In my video I rounded the active parameters to 3 billion and the training data to 24 trillion tokens, and the official figures are 3.46 billion and about 23.6 trillion.
The context window really does reach about 1 million tokens, but Aleph Alpha recommends staying at or under 262,144 for speed and complex tasks.
🔥 Want the exact local models like Kolibri setup I use?
Inside the AI Profit Boardroom I've got the full step-by-step video tutorials, the installable Agent OS, weekly coaching calls and 3,400+ members building real automations.
Why Aleph Alpha built Kolibri AI
Aleph Alpha positions Kolibri as a sovereign open-weight model for regulated work.
Its launch blog says it was built with the EU AI Act, the General-Purpose AI Code of Practice and GDPR in mind from the ground up.
The model card confirms Aleph Alpha is a signatory of the EU GPAI Code of Practice.
The customers it names are public administration, industrial companies and aerospace.
That explains the design choices, from the German-first tokenizer to the decision to go deep in two languages instead of shallow in fifty.
Is Kolibri AI on the GoldieBench leaderboard?
No, Kolibri AI is not on GoldieBench as of 7 October 2026.
GoldieBench gives every model the same one-shot prompts to build a game, a page, a simulation or a visual, then scores the rendered result from 0 to 10.
To bench Kolibri properly I would need to run all of those tasks on hardware that can hold the model, and the official build needs about 78 GB of memory.
I am not going to invent a score for it, so until it runs the full task set there is no GoldieBench number for Kolibri.
What I can do is put Aleph Alpha's own numbers next to the open and local models that have been through the bench.
Kolibri AI benchmarks, as reported by Aleph Alpha
Every number in this table comes from Aleph Alpha's model card, run on their own evaluation framework with Kolibri at high reasoning effort.
That makes these vendor-reported scores, not GoldieBench scores.
| Vendor-reported | Kolibri | Qwen3.5 35B-A3B | Qwen3.6 35B-A3B | Gemma 4 26B-A4B | Nemotron 3 Super | Mistral Small 4 | GPT-OSS 120B | Qwen 3.8 27B (dense) |
|---|---|---|---|---|---|---|---|---|
| Overall English | 75.5 | 74.7 | 71.4 | 71.9 | 73.0 | 63.1 | 72.3 | 80.2 |
| Overall German | 70.8 | 69.8 | 67.3 | 66.3 | 67.9 | 61.4 | 70.2 | 79.9 |
| Agentic average | 63.4 | 63.4 | 62.1 | 54.6 | 54.9 | 40.7 | 54.0 | 66.7 |
| Maths average | 96.5 | 90.1 | 87.8 | 87.4 | 91.1 | 81.4 | 90.7 | 97.8 |
| Code average | 89.3 | 85.0 | 87.7 | 89.0 | 88.3 | 82.0 | 90.8 | 94.2 |
| SWE-Bench Verified | 66.4 | 71.6 | 73.8 | 57.8 | 60.2 | 60.8 | Not reported | 72.6 |
Three things jump out of that table.
First, Kolibri has the best German overall score of the mixture-of-experts models Aleph Alpha tested, which matches what I said in the video.
Second, on the multi-step tool-calling average it beats Qwen 3.6, Nemotron 3 Super and Mistral Small 4, and it ties Qwen3.5 35B-A3B.
Third, the dense Qwen 3.8 27B beats Kolibri on almost every average, including both overall scores.
A commenter under my video said exactly that, and Aleph Alpha's own table agrees with them.
The trade-off is compute, because Qwen 3.8 27B uses all 27 billion parameters on every token while Kolibri uses about 3.46 billion.
There are weak spots too, because Kolibri scores 34.0 on the RGB fact-check test and 51.0 on RGB closed-book questions, the lowest closed-book score in the table.
How Kolibri AI compares with the open and local models on GoldieBench
GoldieBench measures something different from Aleph Alpha's tables, because it scores what a model actually builds in one shot.
So this comparison is about context, not a head-to-head score.
These are the live averages for the open and local models closest to Kolibri.
| Model on GoldieBench | Avg score | Tasks scored | Why it is relevant to Kolibri |
|---|---|---|---|
| GLM-5.2 | 7.77 | 47 | One of the strongest open-weights models on the board |
| Qwable 5 27B Coder | 7.14 | 41 | A fine-tune of Qwen3.6-27B, and the best strictly local builder we have tested |
| Qwen 3.7 | 7.00 | 47 | Alibaba's open-weights Qwen release on the board |
| Agents-A1 | 4.83 | 45 | A 35B mixture of experts with about 3B active, the closest shape to Kolibri on the board |
| Gemma 4 12B MLX | 3.98 | 42 | The fast, small local engine for lightweight jobs |
| Laguna XS 2.1 | 3.93 | 42 | Another small local option |
| Qwythos 9B | 2.98 | 42 | The smallest local model on the board |
One naming trap matters here.
The model called Qwen 3.8 on GoldieBench scores 8.10, but that is Alibaba's huge Qwen3.8-Max-Preview benched through Qoder, not the 27B model in Aleph Alpha's table.
Nemotron and LFM 2.5 are not on the board either, so I am not quoting a GoldieBench score for them.
The full local table, with every demo clickable, is on the local models board.
What the board suggests about Kolibri's shape
Agents-A1 is the most useful reference point, because it is also a mixture of experts with only about 3 billion active parameters.
It averages 4.83 on GoldieBench, well behind the dense Qwable 5 27B at 7.14.
That pattern lines up with Aleph Alpha's own table, where the dense Qwen 3.8 27B beats every mixture-of-experts model including Kolibri.
Small active parameter counts make a model fast and cheap per token, but on one-shot builds the denser models have tended to score higher on our board.
Kolibri is much bigger than Agents-A1, with 78 billion total parameters against 35 billion, so I would not assume it lands in the same place.
The only honest answer is to run it through the tasks, and until then this is context rather than a prediction.
What Kolibri AI needs to run
This is the part that decides whether you can test it at all.
| Build | Size | Memory you need | Status |
|---|---|---|---|
| Official FP8 | About 78 GB | 2× A100 80 GB, 2× H100, 1× H200, 1× B200 or 1× B300 minimum | Official, via vLLM and Aleph Alpha's plugin |
| Community MLX 4-bit | About 41 GiB | A Mac with 64 GB or more | Unofficial, ships its own launcher |
| Community MLX 2-bit | About 24 GiB | A Mac with 36 GB or more | Unofficial, with a bigger quality hit |
| Community GGUF Q4_K_M | About 47.5 GB | Lots of RAM and a patched llama.cpp | Unofficial |
For comparison, Agents-A1's official Q4 GGUF is 21 GB and runs on a 36 GB Mac, and Qwable 5 27B is a 15 GB MLX download.
A viewer said a 78B model needs roughly 100 GB of memory to run properly, and for the official build with room for context that is a fair estimate.
Another asked about a 2 GB graphics card, and the answer is no.
In my video I said you can run it in LM Studio, but as of 7 October 2026 stock LM Studio and Ollama cannot load Kolibri, because llama.cpp, mlx-lm and Ollama support requests are still open.
The working routes today are vLLM with pip install 'aleph-alpha-inference>=1', or a community MLX or GGUF build with its own launcher or patch.
Kolibri AI speed: what we know
I have not measured Kolibri's speed myself, so here is only what the community converters publish.
One MLX converter reports around 52 to 56 tokens per second on an M1 Max.
One GGUF converter reports around 13 to 15 tokens per second for the 4-bit file on a CPU-only desktop with 128 GB of RAM.
Those are their numbers on their machines, so measure it on yours before you build anything around it.
Kolibri AI with Hermes Agent
Kolibri supports reasoning effort levels of none, low, medium and high, plus Hermes-style tool calling on the official server.
That makes it a natural candidate for a local Hermes Agent brain, which is the pairing I showed in the video.
I run a Mac Studio and I don't run much local AI myself, and the best local Hermes model I have tested so far is LFM 2.5 at 2.6 billion parameters because it is crazy fast.
For a full ranking of the local models we have benched for Hermes, read the best local model for Hermes Agent.
My advice is to treat Kolibri as the high-quality German lane in your stack, and keep a small fast model for the high-volume jobs.
Vendor benchmarks versus GoldieBench: why both matter
Aleph Alpha's tables and GoldieBench answer different questions, and it helps to know which one you are reading.
Aleph Alpha's tables use standard tests like GPQA Diamond, AIME, SWE-Bench Verified and tool-calling suites, scored automatically.
Those tests are great for knowledge, maths and multi-turn tool use, and they are where Kolibri looks strongest.
GoldieBench asks a model to build a complete thing in one shot, renders the result and scores what you would actually see on screen.
That rewards a model that can hold a whole project in its head and write clean front-end code without a second try.
A model can be excellent at German document questions and only average at one-shot game builds, and that would not make it a bad model.
It would just mean you should use it for the job it was built for.
Kolibri was built for reasoning, retrieval over your own documents, tool calling, coding and German and English assistants.
So if your work is German contracts, long reports or private agent workflows, its vendor scores are the more relevant signal.
If your work is shipping one-shot builds, the models with real GoldieBench scores are the safer bet today.
How to test Kolibri AI yourself, the GoldieBench way
You do not need my whole bench to get a useful answer for your own work.
First, pick five real tasks from your week, such as a German email reply, a contract summary, a small script and a tool-calling job.
Second, run each task once on Kolibri with reasoning effort set to medium, and save the output without editing it.
Third, run the same five tasks once on the model you use today, such as Qwable 5 27B or Agents-A1 if you work locally.
Fourth, score every output from 0 to 10 on whether you could ship it as it is.
Fifth, write down the time each one took, because a slower model has to be clearly better to earn its place.
That is the same one-shot, score-what-you-see approach GoldieBench uses, scaled down to your own jobs.
If Kolibri wins on your German or long-document tasks, give it that lane and keep a smaller model for the rest.
Who Kolibri AI is for
Kolibri is for teams working in German and English who want an open model they fully control.
It suits European businesses with strict data rules, because it runs on your own hardware under Apache 2.0.
It suits anyone with a GPU server or a 64 GB-plus Mac who wants long-context document work and tool calling.
It is not for a 16 GB laptop, where the smaller models on the local board will serve you far better.
Also On Our Network
🌐 the operator's step-by-step Kolibri AI setup
🌐 seven real business uses for Kolibri AI
🌐 our scored Kolibri AI review
🌐 running Kolibri AI inside an Agent OS
🌐 Julian's guide to running Hermes free forever on a local model
Until it runs the full task set, the honest summary is that Kolibri AI looks strong on its maker's own tests and still has to prove itself on GoldieBench.
FAQ
Is Kolibri AI on GoldieBench?
No. As of 7 October 2026 Kolibri has not been run through the GoldieBench tasks, so there is no GoldieBench score for it. Every Kolibri benchmark quoted here is vendor-reported by Aleph Alpha.
Is Kolibri AI better than Qwen?
On Aleph Alpha's own table, Kolibri beats Qwen3.6 35B-A3B on overall English and German scores and ties Qwen3.5 35B-A3B on the agentic average. The dense Qwen 3.8 27B scores higher than Kolibri overall, with 80.2 in English and 79.9 in German against 75.5 and 70.8.
What is the closest model to Kolibri on GoldieBench?
Agents-A1 is the closest in shape, because it is also a mixture of experts with about 3 billion active parameters. It averages 4.83 on GoldieBench across 45 tasks, while the dense Qwable 5 27B averages 7.14.
How much memory does Kolibri AI need?
The official FP8 weights are about 78 GB, with two 80 GB A100s or one H200 as Aleph Alpha's minimum. Community 4-bit MLX builds need a 64 GB Mac, and the smallest community builds need about 36 GB.
Is it Kolibri or Colibri?
The official name is Kolibri, the German word for hummingbird. Colibri AI is a common alternative spelling for the same model.
About Julian
I'm Julian Goldie — AI entrepreneur, SEO expert, and founder of the AI Profit Boardroom (3,400+ members).
I help business owners scale with AI agents, automation, and SEO.
400K+ YouTube subscribers. 7-figure AI agency (Goldie Agency). Daily training inside the Boardroom.
Related reading
→ The Best Local Model For Hermes Agent — Decided By 45 Real Builds
→ The Best LLMs For Hermes Agent, Ranked By Real Work
→ The Best Free AI Model For Hermes Agent (All Three $0 Lanes)
🌐 Sister-site take: read this on agentos.guide
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 3,400+ founders shipping with it every day all live inside the AI Profit Boardroom.
📺 Video notes + links to the tools 👉 AI Profit Boardroom
🆓 Get a FREE AI Course + Community + 1,000 AI Agents 👉 AI Money Lab
🎥 Learn how I make these videos 👉 agentos.guide