⭐ Get the Agent OS + join 3,400+ founders inside the AI Profit Boardroom → Join AIPB ($69/mo)

Kolibri AI: What It Is, How It Benchmarks, And How It Compares With The Local Models On GoldieBench

By Julian Goldie · 2026-10-07 · GoldieBench Blog

Kolibri AI: What It Is, How It Benchmarks, And How It Compares With The Local Models On GoldieBench — illustrated hero

Kolibri AI is a free, open-weight model from the German company Aleph Alpha, released on 3 October 2026 with 78.1 billion parameters and about 3.46 billion active per token.

People also search for it as Colibri, but the official name is Kolibri, which is German for hummingbird.

The first question everyone asks me is simple: is it any good compared with the models we already know?

I run GoldieBench, where 40 models have now produced 1,609 one-shot demos across 50 tasks.

Kolibri is not on the board yet, and I'll say that plainly up front.

So in this post I'll show you what Kolibri is, what Aleph Alpha's own benchmarks say, what it really takes to run, and how it lines up against the local and open models that are on the board.

What Kolibri AI is

Kolibri AI is Kolibri 1, Aleph Alpha's mixture-of-experts reasoning model with a focus on German and English.

Aleph Alpha is an AI company based in Heidelberg, and the model lives on Hugging Face as Aleph-Alpha/Kolibri-1.

Here is the spec sheet from the official model card.

SpecOfficial detail
Released3 October 2026
LicenceApache 2.0, and the repo is not gated
Total parameters78.1 billion
Active per tokenAbout 3.46 billion
Experts384 per layer, with 6 routed and 1 shared, across 50 layers
LanguagesGerman and English, both native
Context262,144 tokens native, validated up to 1,048,576
TrainingAbout 20 trillion pre-training tokens plus 3.44 trillion mid-training and 201 billion long-context tokens
Knowledge cutoff18 June 2026
FeaturesReasoning mode with effort levels, plus tool calling
Input and outputText only

In my video I rounded the active parameters to 3 billion and the training data to 24 trillion tokens, and the official figures are 3.46 billion and about 23.6 trillion.

The context window really does reach about 1 million tokens, but Aleph Alpha recommends staying at or under 262,144 for speed and complex tasks.

🔥 Want the exact local models like Kolibri setup I use?

Inside the AI Profit Boardroom I've got the full step-by-step video tutorials, the installable Agent OS, weekly coaching calls and 3,400+ members building real automations.

→ Get access here

Why Aleph Alpha built Kolibri AI

Aleph Alpha positions Kolibri as a sovereign open-weight model for regulated work.

Its launch blog says it was built with the EU AI Act, the General-Purpose AI Code of Practice and GDPR in mind from the ground up.

The model card confirms Aleph Alpha is a signatory of the EU GPAI Code of Practice.

The customers it names are public administration, industrial companies and aerospace.

That explains the design choices, from the German-first tokenizer to the decision to go deep in two languages instead of shallow in fifty.

Is Kolibri AI on the GoldieBench leaderboard?

No, Kolibri AI is not on GoldieBench as of 7 October 2026.

GoldieBench gives every model the same one-shot prompts to build a game, a page, a simulation or a visual, then scores the rendered result from 0 to 10.

To bench Kolibri properly I would need to run all of those tasks on hardware that can hold the model, and the official build needs about 78 GB of memory.

I am not going to invent a score for it, so until it runs the full task set there is no GoldieBench number for Kolibri.

What I can do is put Aleph Alpha's own numbers next to the open and local models that have been through the bench.

Kolibri AI benchmarks, as reported by Aleph Alpha

Every number in this table comes from Aleph Alpha's model card, run on their own evaluation framework with Kolibri at high reasoning effort.

That makes these vendor-reported scores, not GoldieBench scores.

Vendor-reportedKolibriQwen3.5 35B-A3BQwen3.6 35B-A3BGemma 4 26B-A4BNemotron 3 SuperMistral Small 4GPT-OSS 120BQwen 3.8 27B (dense)
Overall English75.574.771.471.973.063.172.380.2
Overall German70.869.867.366.367.961.470.279.9
Agentic average63.463.462.154.654.940.754.066.7
Maths average96.590.187.887.491.181.490.797.8
Code average89.385.087.789.088.382.090.894.2
SWE-Bench Verified66.471.673.857.860.260.8Not reported72.6

Three things jump out of that table.

First, Kolibri has the best German overall score of the mixture-of-experts models Aleph Alpha tested, which matches what I said in the video.

Second, on the multi-step tool-calling average it beats Qwen 3.6, Nemotron 3 Super and Mistral Small 4, and it ties Qwen3.5 35B-A3B.

Third, the dense Qwen 3.8 27B beats Kolibri on almost every average, including both overall scores.

A commenter under my video said exactly that, and Aleph Alpha's own table agrees with them.

The trade-off is compute, because Qwen 3.8 27B uses all 27 billion parameters on every token while Kolibri uses about 3.46 billion.

There are weak spots too, because Kolibri scores 34.0 on the RGB fact-check test and 51.0 on RGB closed-book questions, the lowest closed-book score in the table.

How Kolibri AI compares with the open and local models on GoldieBench

GoldieBench measures something different from Aleph Alpha's tables, because it scores what a model actually builds in one shot.

So this comparison is about context, not a head-to-head score.

These are the live averages for the open and local models closest to Kolibri.

Model on GoldieBenchAvg scoreTasks scoredWhy it is relevant to Kolibri
GLM-5.27.7747One of the strongest open-weights models on the board
Qwable 5 27B Coder7.1441A fine-tune of Qwen3.6-27B, and the best strictly local builder we have tested
Qwen 3.77.0047Alibaba's open-weights Qwen release on the board
Agents-A14.8345A 35B mixture of experts with about 3B active, the closest shape to Kolibri on the board
Gemma 4 12B MLX3.9842The fast, small local engine for lightweight jobs
Laguna XS 2.13.9342Another small local option
Qwythos 9B2.9842The smallest local model on the board

One naming trap matters here.

The model called Qwen 3.8 on GoldieBench scores 8.10, but that is Alibaba's huge Qwen3.8-Max-Preview benched through Qoder, not the 27B model in Aleph Alpha's table.

Nemotron and LFM 2.5 are not on the board either, so I am not quoting a GoldieBench score for them.

The full local table, with every demo clickable, is on the local models board.

What the board suggests about Kolibri's shape

Agents-A1 is the most useful reference point, because it is also a mixture of experts with only about 3 billion active parameters.

It averages 4.83 on GoldieBench, well behind the dense Qwable 5 27B at 7.14.

That pattern lines up with Aleph Alpha's own table, where the dense Qwen 3.8 27B beats every mixture-of-experts model including Kolibri.

Small active parameter counts make a model fast and cheap per token, but on one-shot builds the denser models have tended to score higher on our board.

Kolibri is much bigger than Agents-A1, with 78 billion total parameters against 35 billion, so I would not assume it lands in the same place.

The only honest answer is to run it through the tasks, and until then this is context rather than a prediction.

What Kolibri AI needs to run

This is the part that decides whether you can test it at all.

BuildSizeMemory you needStatus
Official FP8About 78 GB2× A100 80 GB, 2× H100, 1× H200, 1× B200 or 1× B300 minimumOfficial, via vLLM and Aleph Alpha's plugin
Community MLX 4-bitAbout 41 GiBA Mac with 64 GB or moreUnofficial, ships its own launcher
Community MLX 2-bitAbout 24 GiBA Mac with 36 GB or moreUnofficial, with a bigger quality hit
Community GGUF Q4_K_MAbout 47.5 GBLots of RAM and a patched llama.cppUnofficial

For comparison, Agents-A1's official Q4 GGUF is 21 GB and runs on a 36 GB Mac, and Qwable 5 27B is a 15 GB MLX download.

A viewer said a 78B model needs roughly 100 GB of memory to run properly, and for the official build with room for context that is a fair estimate.

Another asked about a 2 GB graphics card, and the answer is no.

In my video I said you can run it in LM Studio, but as of 7 October 2026 stock LM Studio and Ollama cannot load Kolibri, because llama.cpp, mlx-lm and Ollama support requests are still open.

The working routes today are vLLM with pip install 'aleph-alpha-inference>=1', or a community MLX or GGUF build with its own launcher or patch.

Kolibri AI speed: what we know

I have not measured Kolibri's speed myself, so here is only what the community converters publish.

One MLX converter reports around 52 to 56 tokens per second on an M1 Max.

One GGUF converter reports around 13 to 15 tokens per second for the 4-bit file on a CPU-only desktop with 128 GB of RAM.

Those are their numbers on their machines, so measure it on yours before you build anything around it.

Kolibri AI with Hermes Agent

Kolibri supports reasoning effort levels of none, low, medium and high, plus Hermes-style tool calling on the official server.

That makes it a natural candidate for a local Hermes Agent brain, which is the pairing I showed in the video.

I run a Mac Studio and I don't run much local AI myself, and the best local Hermes model I have tested so far is LFM 2.5 at 2.6 billion parameters because it is crazy fast.

For a full ranking of the local models we have benched for Hermes, read the best local model for Hermes Agent.

My advice is to treat Kolibri as the high-quality German lane in your stack, and keep a small fast model for the high-volume jobs.

Vendor benchmarks versus GoldieBench: why both matter

Aleph Alpha's tables and GoldieBench answer different questions, and it helps to know which one you are reading.

Aleph Alpha's tables use standard tests like GPQA Diamond, AIME, SWE-Bench Verified and tool-calling suites, scored automatically.

Those tests are great for knowledge, maths and multi-turn tool use, and they are where Kolibri looks strongest.

GoldieBench asks a model to build a complete thing in one shot, renders the result and scores what you would actually see on screen.

That rewards a model that can hold a whole project in its head and write clean front-end code without a second try.

A model can be excellent at German document questions and only average at one-shot game builds, and that would not make it a bad model.

It would just mean you should use it for the job it was built for.

Kolibri was built for reasoning, retrieval over your own documents, tool calling, coding and German and English assistants.

So if your work is German contracts, long reports or private agent workflows, its vendor scores are the more relevant signal.

If your work is shipping one-shot builds, the models with real GoldieBench scores are the safer bet today.

How to test Kolibri AI yourself, the GoldieBench way

You do not need my whole bench to get a useful answer for your own work.

First, pick five real tasks from your week, such as a German email reply, a contract summary, a small script and a tool-calling job.

Second, run each task once on Kolibri with reasoning effort set to medium, and save the output without editing it.

Third, run the same five tasks once on the model you use today, such as Qwable 5 27B or Agents-A1 if you work locally.

Fourth, score every output from 0 to 10 on whether you could ship it as it is.

Fifth, write down the time each one took, because a slower model has to be clearly better to earn its place.

That is the same one-shot, score-what-you-see approach GoldieBench uses, scaled down to your own jobs.

If Kolibri wins on your German or long-document tasks, give it that lane and keep a smaller model for the rest.

Who Kolibri AI is for

Kolibri is for teams working in German and English who want an open model they fully control.

It suits European businesses with strict data rules, because it runs on your own hardware under Apache 2.0.

It suits anyone with a GPU server or a 64 GB-plus Mac who wants long-context document work and tool calling.

It is not for a 16 GB laptop, where the smaller models on the local board will serve you far better.

Also On Our Network

🌐 the operator's step-by-step Kolibri AI setup

🌐 seven real business uses for Kolibri AI

🌐 our scored Kolibri AI review

🌐 running Kolibri AI inside an Agent OS

🌐 Julian's guide to running Hermes free forever on a local model

Until it runs the full task set, the honest summary is that Kolibri AI looks strong on its maker's own tests and still has to prove itself on GoldieBench.

FAQ

Is Kolibri AI on GoldieBench?

No. As of 7 October 2026 Kolibri has not been run through the GoldieBench tasks, so there is no GoldieBench score for it. Every Kolibri benchmark quoted here is vendor-reported by Aleph Alpha.

Is Kolibri AI better than Qwen?

On Aleph Alpha's own table, Kolibri beats Qwen3.6 35B-A3B on overall English and German scores and ties Qwen3.5 35B-A3B on the agentic average. The dense Qwen 3.8 27B scores higher than Kolibri overall, with 80.2 in English and 79.9 in German against 75.5 and 70.8.

What is the closest model to Kolibri on GoldieBench?

Agents-A1 is the closest in shape, because it is also a mixture of experts with about 3 billion active parameters. It averages 4.83 on GoldieBench across 45 tasks, while the dense Qwable 5 27B averages 7.14.

How much memory does Kolibri AI need?

The official FP8 weights are about 78 GB, with two 80 GB A100s or one H200 as Aleph Alpha's minimum. Community 4-bit MLX builds need a 64 GB Mac, and the smallest community builds need about 36 GB.

Is it Kolibri or Colibri?

The official name is Kolibri, the German word for hummingbird. Colibri AI is a common alternative spelling for the same model.

About Julian

I'm Julian Goldie — AI entrepreneur, SEO expert, and founder of the AI Profit Boardroom (3,400+ members).

I help business owners scale with AI agents, automation, and SEO.

400K+ YouTube subscribers. 7-figure AI agency (Goldie Agency). Daily training inside the Boardroom.

→ Get my best AI training inside the AI Profit Boardroom

Related reading

→ The Best Local Model For Hermes Agent — Decided By 45 Real Builds

→ The Best LLMs For Hermes Agent, Ranked By Real Work

→ The Best Free AI Model For Hermes Agent (All Three $0 Lanes)

🌐 Sister-site take: read this on agentos.guide

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 3,400+ founders shipping with it every day all live inside the AI Profit Boardroom.

3,400+founders
258documented wins
38countries
$69/momonthly

📺 Video notes + links to the tools 👉 AI Profit Boardroom

🆓 Get a FREE AI Course + Community + 1,000 AI Agents 👉 AI Money Lab

🎥 Learn how I make these videos 👉 agentos.guide

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 3,400+ founders shipping with it every day all live inside the AI Profit Boardroom.

3,400+founders
258documented wins
38countries
$69/momonthly