⭐ Get the Agent OS + join 3,400+ founders inside the AI Profit Boardroom → Join AIPB ($69/mo)

GoldieBench Blog · 9 min read

The Best Hermes Agent Model In 2026: Five Category Winners, With The Live Scores Behind Them

The best Hermes Agent model is Claude Opus 5 at 8.27. See the best free, local, cheap and long-loop picks with live GoldieBench scores for each one.

The Best Hermes Agent Model In 2026: Five Category Winners, With The Live Scores Behind Them — illustrated hero

The best Hermes Agent model in 2026 is Claude Opus 5, which averages 8.27 on GoldieBench and runs inside Hermes on a normal Claude subscription.

My best free pick is Solar Mini 4, my best local pick is Agents-A1, my best cheap pick is MiniMax M3, and my best pick for long agent loops is Kimi K3.

Hermes Agent is an open-source AI agent from Nous Research, and it doesn't ship with a brain of its own.

The model is that brain, so it reads every prompt Hermes builds, picks the next tool to call and decides when the job is finished.

Hermes supplies the tools, the memory and the safety checks, which means you can swap the model without rebuilding anything.

I'll give you the category winners first, then the ranked table with live scores, and then what those scores can and can't tell you about agent work.

The five category winners at a glance

  • 🥇 Best overall: Claude Opus 5 is the top single model on the board at 8.27.
  • 🆓 Best free: Solar Mini 4 is free on Nous Portal for a limited time, and it isn't on the board yet.
  • 💻 Best local: Agents-A1 averages 4.83 and is tuned for tool calling.
  • 💸 Best cheap: MiniMax M3 averages 7.97 at $0.30 per million input tokens.
  • 🔁 Best for long agent loops: Kimi K3 averages 7.89 and holds about one million tokens of context.

How I ranked the best Hermes Agent model picks

This ranking is my own, and it's built from the models I've actually wired into Hermes and used.

I weighed four things.

The first is build quality, using the live GoldieBench average for every model that has one.

The second is cost, including whether a subscription you already own can cover it.

The third is how easy the model is to connect to Hermes.

The fourth is how the model behaves when the agent runs many steps in a row.

The one thing these scores do not measure

GoldieBench scores one-shot builds.

Every model gets the same prompt and one attempt, the result is rendered, and it's scored from 0 to 10.

That's a strong test of how well a model builds in a single pass.

It is not a test of agent loops.

An agent loop is a long chain of tool calls where the model must remember what it did twenty steps ago.

So a model can score well here and still drift on a long job, and a model can score low here and still be a good tool caller.

I'll tell you each time a pick rests on hands-on use and not on a score.

The best Hermes Agent models, ranked with live scores

RankModelMakerLive averageTasks scoredGold medalsPrice on the board
1Claude Opus 5Anthropic8.275013$5 in, $25 out per million
2GPT-5.6 SolOpenAI8.16502$5 in, $30 out per million
3Kimi K3Moonshot AI7.89508$3 per million input
4MiniMax M3MiniMax7.97472$0.30 in, $1.50 out per million
5GLM-5.2Zhipu7.77475Open weights, free for individuals
6Grok 4.7xAI7.15200$2 in, $6 out per million
7Solar Mini 4Upstage AINot on the boardNot scoredNot scoredFree on Nous Portal for a limited time
8Agents-A1InternScience4.83450Free, runs locally

The order is my ranking for agent use, which is why Kimi K3 sits above MiniMax M3 with a slightly lower average.

Every number in that table is a live average on 10 October 2026.

Best overall: Claude Opus 5

Claude Opus 5 is Anthropic's flagship, and it has the highest average of any single model on the board.

It also has 13 gold medals, which is more than any other model in this ranking.

The reason it's my number one for Hermes is the route as much as the score.

An official Nous Research plugin, announced on 22 September 2026, lets Hermes use the Claude Code login you already have.

You don't need an API key, and Hermes still controls the tools, the memory and the approvals.

The plugin needs Hermes 0.21.4 or newer.

One setting chooses the Claude you get, and it accepts opus, sonnet, haiku or fable.

Claude Sonnet 5 averages 7.01, so it's a fair daily driver on the same plugin. Claude Opus 5.5 averages 7.57, which is lower than Opus 5 on one-shot builds.

That's a useful reminder that a newer model doesn't always score higher on this board.

Runner-up: GPT-5.6 Sol

GPT-5.6 Sol averages 8.16 across 50 scored tasks.

It has two golds, nine silvers and ten bronzes, so it lands on the podium often.

I run GPT models in Hermes through an OpenRouter profile.

It's the brain I'd choose as a second opinion on important work.

Best for long agent loops: Kimi K3

Kimi K3 averages 7.89 with eight gold medals.

It holds about one million tokens of context, and Moonshot AI tuned it for long-horizon agent work.

That design is why I pick it for long loops, and this pick rests on hands-on use because the board doesn't test loops.

K3 is included in the Kimi coding plan, which is a flat subscription, so a long job doesn't run up a per-token bill.

DeepSeek is my alternative for this job.

DeepSeek V4 Flash is cheap and was retrained for agent loops, but it's currently unranked on the board, so I can't quote a score for it.

Best cheap: MiniMax M3

MiniMax M3 averages 7.97 across 47 scored tasks.

It's the cheapest big-context model on the board, at $0.30 per million input tokens and $1.50 per million output tokens.

That average is higher than Kimi K3 and GLM-5.2, at a fraction of the flagship price.

I use it with Hermes for automations that call a lot of tools.

A cheap model changes how you use an agent, because you stop asking whether a job is worth the tokens.

You let it watch a folder, summarise every session or run on a schedule.

GLM-5.2 is my other value pick at 7.77, with five golds and open weights.

GLM-5.3 is live on the GLM Coding Plan, but it isn't scored on the board yet.

GLM sometimes cuts long outputs short, so I keep a fallback behind it.

Fast and cheap: Grok 4.7

Grok 4.7 averages 7.15, but only across 20 scored tasks, so treat it as a smaller sample.

The older Grok entry averages 8.09 across 43 scored tasks.

In my own Hermes test, Grok 4.7 fixed and tested a small bug in 13.9 seconds.

I run it as a profile through OpenRouter.

Best free: Solar Mini 4

Solar Mini 4 is a mixture-of-experts model from Upstage AI in South Korea.

It has 3 billion active parameters out of 35 billion and a context window of 500,000 tokens.

It's free on Nous Portal for a limited time, and you pick the variant marked free in the Hermes model list.

Solar Mini 4 isn't on GoldieBench, so I have no score to show you.

In my hands-on test it replied fast and handled a news research task with sources.

I asked it for AI automation news from the last seven days, and it came back with a breakdown, sources and a note on what stood out.

It even reported a launch I hadn't heard about.

It isn't frontier level, and free hosted models can be rate limited.

My full test is in Solar Mini 4 with Hermes.

If you want a free model that is on the board, Laguna XS 2.1 averages 3.93 and has a free tier on OpenRouter.

Best local: Agents-A1

Agents-A1 is an open-weight model from InternScience that's tuned for long-horizon tool work.

It averages 4.83, which is low for one-shot builds.

It runs at about 95 tokens a second on my 36GB Mac, and that speed is what keeps a local agent loop usable.

This is the clearest case where a build score and agent usefulness point in different directions.

Local modelLive averageTasks scoredWhat I use it for
Qwable 5 27B Coder7.1441The best local build quality, but slow for loops
Agents-A14.8345Local agent loops and tool calling
Gemma 4 12B MLX3.9842A light helper that runs on 16GB
Laguna XS 2.13.9342A small agentic coder

LFM2.5 from LiquidAI is another local option that fits on an 8GB laptop, but it isn't on the board.

The local models board has every local score with a playable demo.

How to connect any of these models to Hermes

You run the hermes model command to choose a provider and a default model.

I keep one Hermes profile per model, so each brain has its own settings.

I switch between them with the profile flag, and I check which model is really running with the status command.

Hermes also has a fallback setting that tries a second provider when the first one fails.

My routing is simple.

A strong subscription model is my main brain, a cheap or flat-rate model handles long and repeat work, and a free or local model handles the small jobs.

Which Hermes Agent model fits your budget

If you already pay for Claude, use Claude Opus 5 as your main brain, because the plugin adds no new bill.

If you already pay for the Kimi coding plan, use Kimi K3, because it's included in the plan.

If you pay per token, start with MiniMax M3, because the price is low enough to let an agent run all day.

If you pay for nothing, start with Solar Mini 4 on Nous Portal and a local model through Ollama.

If your data can't leave your machine, use a local model and keep the tasks simple.

Most people end up with two or three brains and not one.

The strong brain handles planning and anything a client will read.

The cheap brain handles the repeat work.

The free or local brain handles summaries, file reads and small tool chains.

Common mistakes when picking a Hermes Agent model

The first mistake is choosing on the average alone and ignoring how many tasks were scored.

Grok 4.7 has 20 scored tasks and Claude Opus 5 has 50, so those two averages don't carry the same weight.

The second mistake is assuming the newest model is the best one, when Opus 5.5 sits below Opus 5 on this board.

The third mistake is picking the paid variant of a free model by accident, which is easy to do with Solar Mini 4.

The fourth mistake is running every small job on a flagship and then rationing the agent because of the bill.

The fifth mistake is trusting what a model says about itself.

When I asked Kimi K3 what it was, it gave me the name of an older model, so I check the status line instead.

The sixth mistake is having no fallback, because one provider outage then stops the whole agent.

How to read a GoldieBench score for agent work

A high average means the model produces a working, good-looking build from one prompt.

Gold medals mean the model was the best on a task, not just good.

The tasks-scored column tells you how much evidence is behind the average.

None of those numbers tell you how a model handles 200 tool calls.

Use the score to shortlist, and then run one real agent job before you commit.

You can put any models side by side on the compare page, and the methodology page explains how every build is scored.

More on Hermes models from this blog

My earlier post on the best LLMs for Hermes Agent covers the three tiers of brains.

My post on the best local model for Hermes Agent goes deeper on the local board.

My post on the best free AI model for Hermes Agent covers the three free lanes.

This post is the one that picks a winner for each category.

The model is only one part of the stack, so read the best Hermes Agent setup and the best Hermes Agent memory next.

Also On Our Network

Shortlist with the live scores, test one real agent job, and you'll find the best Hermes Agent model for your own work.

FAQ

What is the best Hermes Agent model?

Claude Opus 5 is the best Hermes Agent model in 2026. It is the top single model on GoldieBench at 8.27 with 13 gold medals, and an official Nous Research plugin runs it inside Hermes on a Claude subscription.

Does GoldieBench measure agent loops?

No. GoldieBench scores one-shot builds, where each model gets one prompt and one attempt. It does not test long tool-calling loops, so use the scores to shortlist and then test a real agent job.

Is Solar Mini 4 on GoldieBench?

No. Solar Mini 4 from Upstage AI is not on the board yet. It is my best free pick from hands-on use, because it is fast and free on Nous Portal for a limited time.

What is the best cheap model for Hermes Agent?

MiniMax M3 is my best cheap pick. It averages 7.97 on GoldieBench and is listed at $0.30 per million input tokens and $1.50 per million output tokens.

Why does Agents-A1 rank as best local with a 4.83 average?

Because the average measures one-shot builds. Agents-A1 is tuned for tool calling and runs at about 95 tokens a second on a 36GB Mac, which matters more for a local agent loop. Qwable 5 27B Coder scores 7.14 for build quality but is slower.

Does Hermes Agent ship its own model?

No. Hermes Agent is the agent, and you choose the model. It works with subscriptions, API keys, free hosted models and local models through Ollama.

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 3,400+ founders shipping with it every day all live inside the AI Profit Boardroom.

3,400+founders
258documented wins
38countries
$69/momonthly