GoldieBench Blog · 10 min read
Hermes Herald AI Agent: Which Brain To Run It On, With Real GoldieBench Scores
The Hermes Herald AI agent runs on whatever model Hermes uses. Here are the real GoldieBench scores for the brains you can pick inside Herald OS.

The Hermes Herald AI agent is Hermes Agent running as the whole interface of Herald OS, a free, open-source, agent-native operating system.
Herald OS was built by the independent developer Luke The Dev, and it is not affiliated with Nous Research, which makes Hermes Agent itself.
It is released under the MIT licence, so you can download it and run it without paying for the software.
The problem it solves is that Hermes normally lives in a terminal, and Herald OS turns it into a desktop you can talk or type to.
In this post I explain what the Hermes Herald AI agent is, how it works, how to set it up, what it can do and where it still falls short.
Then I finish with which brain to run behind it, using real scores from my own leaderboard.
What the Hermes Herald AI agent is
Herald OS calls itself "an agent-native operating system, with Hermes Agent as the interface".
You talk or type to Hermes, and it works across the whole machine, opening apps, organising files, running routines and building software.
It is MIT licensed and free, and it is alpha, version 0.1, with its first release on 6 October 2026.
It is not affiliated with Nous Research, which makes Hermes Agent itself.
How the Hermes Herald AI agent works
Herald OS does not fork Hermes.
It starts your normal Hermes install in the background with hermes serve and adds a Hermes plugin called herald-os-bridge.
That plugin gives Hermes typed system tools for files, apps, processes, logs and the Herald interface, each with a permission tier.
Read tools run straight away, file changes ask first, and trashing files or stopping programs asks every time.
Some locations, such as your SSH keys, your keychains and Hermes's own credentials, are always refused, and Settings shows an audit log of what the agent did.
Because it is your normal Hermes underneath, your skills, memory and scheduled jobs come with you.
Herald OS is not the Hermes "Herald Release"
Some people searching this phrase mean the Hermes update, not the operating system.
Nous Research codenamed Hermes Agent v0.20.0 "The Herald Release", and it shipped on 3 August 2026.
That update added conversational voice with barge-in, A2A agent-to-agent support, signed outbound webhooks and grounded citations.
Herald OS is a separate, independent project by Luke The Dev, and its from-source route asks for Hermes 0.21.3 or later.
The model advice further down applies to both, because it is the same Hermes making the calls.
How to set up the Hermes Herald AI agent
First, install Hermes Agent with the official install script and run hermes setup to sign in to a model provider.
Second, download the Herald OS DMG for Apple Silicon from the GitHub releases page and drag it into Applications.
Third, run xattr -dr com.apple.quarantine "/Applications/Herald OS.app" once, because the alpha build is not notarized yet.
Fourth, open Herald OS, and it starts your Hermes in the background and links the bridge plugin on its first run.
Cmd+Ctrl+F toggles fullscreen, and Cmd+Q quits and stops the backend it started.
I had Claude set the whole thing up for me in about two minutes, and my existing Hermes automations synced across straight away.
Linux, Arch, Omarchy and virtual machine routes are in my Hermes Agent OS GitHub install guide, and the full screen-by-screen tour is in the Herald OS explainer.
What the Hermes Herald AI agent can do
The Overview greets you with "Good morning, Julian. Your day already in motion." and shows today's schedule from your scheduled tasks.
"Pick up where you left off" suggests up to three threads from the names of your recent documents, projects, chats and calendar events, and it is off until you switch it on.
Spaces split your work, and my screen showed Ideas, Work and Personal.
Memory shows everything Hermes remembers, and it showed my Telegram details on camera, so check it before you screen-share.
Files lists recent, favourites, projects, downloads, shared and deleted items.
Automations synced across Herald OS, Hermes Desktop and my Agent OS, because Herald uses Hermes's own scheduled jobs.
Connections is meant to hold your tools, apps and external accounts.
Settings covers autonomy, where agents run tasks including Docker, background agents, activity history, voice, usage, software and plugins.
Voice uses Alt+Space on a Mac, with a free engine by default and an opt-in Live engine that the docs price at $0.05 a minute.
Honest limits
Herald OS is alpha, and the interface was cut off slightly at the top on my screen.
Connections did not work or sync for me, so do not build anything important on them yet.
The Mac build is not notarized yet, so only download it from the official releases page.
And keep backups of anything you let an agent touch, which the Herald OS README says too.
Which brain to run behind the Hermes Herald AI agent
Herald OS does not ship its own model, because the Hermes Herald AI agent uses whatever brain your Hermes install is set up with.
That makes the model the single biggest choice you make once the agent is running.
To compare the options, I use GoldieBench, the leaderboard where I score models on real one-shot builds.
How the Hermes Herald AI agent chooses its brain
There are three places the model gets set.
| Where | What you do | When to use it |
|---|---|---|
hermes setup | You pick a provider and sign in the first time you install Hermes. | Day one, before you open Herald OS. |
| Settings, Preferred model | You choose the model Hermes uses for most tasks from inside Herald OS. | Switching brains without leaving the desktop. |
hermes model | You change the model from the terminal. | Scripting a switch or fixing a broken provider. |
| Settings, Software | You install Ollama or LM Studio, and once a model is loaded, Use with Hermes points Hermes at it. | Running a free local brain. |
The providers Herald OS lists are a Nous Portal account or an API key for OpenRouter, OpenAI, Anthropic, a local model and others.
My own settings showed a model labelled "GPT-6 Astra", plugged in through a CLI.
That model is not on GoldieBench, so I will not give it a score it has not earned.
Herald OS itself is free, but the model is billed by its provider unless you pick a free one.
The Usage page shows the tokens Hermes used this week and this month and what is left on your plan, with a warning at 90%.
What GoldieBench measures, and why it matters here
GoldieBench gives every model the same one-shot prompt to build a game, a page, a simulation or a visual, then scores the rendered result from 0 to 10.
That maps neatly onto the build side of the Hermes Herald AI agent.
When you start a Mission like "build out a new website", or ask the Studio to "build a website for a hair salon", the brain behind Hermes is doing exactly the kind of one-shot build the board scores.
It does not measure tool calling, approvals or how politely the agent asks before moving a file.
So read these scores as "how good is this brain at building things", and judge the agent side in your own Herald OS setup.
The brains for the Hermes Herald AI agent, scored
These are the live numbers from the board as of 8 October 2026.
| Brain | GoldieBench avg | Tasks scored | What it is good for in Herald OS |
|---|---|---|---|
| Fusion (OpenRouter) | 8.59 | 47 | The top of the whole board, as a multi-model panel. |
| Claude Opus 5 | 8.27 | 50 | The highest single model, for Missions and Studio builds that must work first time. |
| Hermes MoA | 8.17 | 47 | Hermes's own Mixture of Agents mode, where a panel drafts and an aggregator merges. |
| GPT-5.6 Sol | 8.16 | 50 | A second frontier choice that sits level with Hermes MoA. |
| MiniMax M3 | 7.97 | 47 | Near-frontier scores at $0.30 per million input tokens. |
| Kimi K3 | 7.89 | 50 | A strong all-rounder at $3 per million input tokens. |
| GLM-5.2 | 7.77 | 47 | Open weights, so you are never locked to one provider. |
| Claude Opus 5.5 | 7.57 | 50 | Anthropic's newest Opus, which sits below Opus 5 on this board. |
| Grok 4.7 | 7.15 | 20 | A big jump on Grok 4.6, which scores 5.95 over the same 20 tasks. |
| Qwable 5 27B Coder | 7.14 | 41 | The best free local builder on the local board. |
| Gemini 3.6 Flash | 7.08 | 50 | A cheap, fast lane at $1.50 per million input tokens. |
| Laguna XS 2.1 | 3.93 | 42 | A free tier, but too weak for builds on this board. |
Meituan's LongCat-2.0 shows 8.12, but it has only 4 tasks scored, so it is provisional.
Which brain I would pick for each part of Herald OS
The Hermes Herald AI agent does very different jobs, so one brain for everything is rarely the smartest move.
| Herald OS job | What the brain does | My pick from the board |
|---|---|---|
| Missions and Studio builds | It plans and builds a site, an app or a report from one goal. | Claude Opus 5 (8.27) or GPT-5.6 Sol (8.16) |
| Automations | It writes a daily 6am recap, an inbox summary or a learning prompt. | MiniMax M3 (7.97) or Gemini 3.6 Flash (7.08) |
| Background agents | It runs several jobs at once, from 1 up to 32 in Settings. | A value model, because the bill multiplies with every agent you add. |
| Voice chats | It answers short spoken questions and runs quick tasks. | A fast lane such as Gemini 3.6 Flash. |
| Private files | It works on things you do not want to leave the machine. | Qwable 5 27B Coder (7.14) through LM Studio or Ollama. |
| Hard research | It needs several opinions merged into one answer. | Hermes MoA (8.17). |
Herald OS only has one Preferred model setting, so in practice this means switching the model for the job, or using separate Hermes profiles.
Why Hermes MoA deserves a look behind Herald OS
Hermes MoA is the one entry on the board that is Hermes-native rather than a single model.
It sends one prompt to a panel of frontier models in parallel, then a named aggregator reads every draft and writes one better final answer.
Its default panel is Claude Opus 4.8 plus GPT-5.5, aggregated by Opus 4.8, all through an OpenRouter key.
That panel scored 8.17 across 47 tasks, which puts it level with GPT-5.6 Sol and just behind Claude Opus 5.
The point is that the system can matter as much as the model.
Every slot in the panel is swappable, so as stronger models land on the board, the panel gets stronger too.
The trade-off is cost and speed, because every answer is several model calls plus an aggregation call.
So I would keep it for research-heavy missions where one wrong answer costs you more than the extra calls.
What background agents do to your bill
Herald OS lets you run anywhere from 1 to 32 background agents at once in Settings.
That is brilliant for getting through a list of jobs, but every one of those agents is calling the same brain.
Here is the input-price gap on the board, using each model's listed price.
| Brain | Listed input price per 1M tokens | GoldieBench avg |
|---|---|---|
| Claude Opus 5 | $5 | 8.27 |
| GPT-5.6 Sol | $5 | 8.16 |
| Kimi K3 | $3 | 7.89 |
| Gemini 3.6 Flash | $1.50 | 7.08 |
| MiniMax M3 | $0.30 | 7.97 |
MiniMax M3's input price is about one-seventeenth of Claude Opus 5's, and it gives up only 0.30 points on the board.
That is why I would never point 10 background agents at a frontier model for routine work.
Put the cheap, strong brain on the busy lanes, and save the frontier brain for the jobs where quality decides the outcome.
The Usage page in Herald OS is where you check that this is actually working.
If you want the longer version of this ranking, the best LLMs for Hermes Agent post goes deeper into the board.
If you want a free brain to test first, the Solar Mini 4 Hermes post covers the free Nous Portal models.
How to switch the brain in the Hermes Herald AI agent
First, open Settings in Herald OS with Cmd+, on a Mac.
Second, go to the Hermes and agents section and open Preferred model.
Third, pick the model you want, or sign in to a new provider from the Account row.
Fourth, check the Usage page after a day, because a frontier brain on a busy automation list adds up fast.
Fifth, if you want a local brain, install LM Studio or Ollama from Settings, load a model and click Use with Hermes.
Because Herald OS uses the same Hermes as Hermes Desktop and my Agent OS, the new model applies everywhere that Hermes profile runs.
GoldieBench scores builds, so a model that tops the board can still fumble a badly worded automation.
Free models can be rate limited, so keep a paid backup in the same profile list.
Also On Our Network
- Agent Operatorsthe step-by-step Hermes Herald AI agent setup for operators
- AI Income Deskhow businesses use the Hermes Herald AI agent day to day
- AI Tool Verdictour feature-by-feature Hermes Herald review with scores
- agentos.guidethe Hermes Herald quick-start inside an Agent OS
- agentos.guideJulian's guide to the Hermes v0.20 Herald Release
Pick a strong brain for missions, a cheap one for automations, and let the board tell you when a new model deserves a place behind the Hermes Herald AI agent.
FAQ
Does Herald OS come with its own model?
No. Herald OS runs your normal Hermes Agent install, so the Hermes Herald AI agent uses whatever model Hermes is set up with. You need a Nous Portal account or an API key for a provider such as OpenRouter, OpenAI or Anthropic, or a local model.
What model does the Hermes Herald AI agent use?
Herald OS uses whatever model your Hermes install is set up with, such as a Nous Portal account or an API key for OpenRouter, OpenAI, Anthropic or a local model. You change it in Settings under Preferred model or with hermes model.
What is the best brain for the Hermes Herald AI agent on GoldieBench?
Claude Opus 5 is the highest single model at 8.27, with GPT-5.6 Sol at 8.16 and Hermes MoA, Hermes's Mixture of Agents mode, at 8.17. The OpenRouter Fusion panel tops the whole board at 8.59.
Is there a free brain that scores well?
Qwable 5 27B Coder runs locally for free and averages 7.14 across 41 tasks. Laguna XS 2.1 has a free tier on OpenRouter but averages only 3.93, so it is weak for builds.
Is GPT-6 Astra on GoldieBench?
No. Julian's Herald OS settings showed a model labelled GPT-6 Astra through a CLI, but it has not been benched on GoldieBench, so it has no score.
Is Herald OS the same as the Hermes Herald Release?
No. The Herald Release is Nous Research's codename for Hermes Agent v0.20.0, released on 3 August 2026. Herald OS is an independent operating system by Luke The Dev that runs on Hermes and is not affiliated with Nous Research.


