Solar Mini 4 Hermes: How To Run It Free, And How It Compares With The Benched Hermes Brains
By Julian Goldie · 2026-10-07 · GoldieBench Blog
Solar Mini 4 Hermes means running Upstage's new Solar Mini 4 model as the brain of Hermes Agent, and right now you can do it for free on Nous Portal.
I run GoldieBench, the leaderboard where 40 models have produced 1,609 one-shot demos across 50 tasks.
So this post does two jobs for you.
First, it shows you exactly how to set up Solar Mini 4 inside Hermes without landing on the paid version.
Second, it uses real numbers from the board to show you which Hermes brains have actually earned their place, and where a model like Solar Mini 4 would fit.
I'll be upfront about one thing straight away.
Solar Mini 4 is not on the GoldieBench leaderboard yet, so I won't give it a score it hasn't earned.
What Solar Mini 4 is, in short
Solar Mini 4 is a compact model from Upstage AI in South Korea.
It's a mixture-of-experts model with 35 billion parameters in total and 3 billion active at a time.
Its context window is roughly 500,000 tokens, and OpenRouter lists it at 524,288.
Artificial Analysis gives it 24 on its Intelligence Index, which it says is one point above Nemotron 3 Ultra with 55 billion active parameters.
Artificial Analysis also describes it as proprietary, so there are no weights to run locally.
It's free for two weeks on Nous Portal, the Nous Research portal that Hermes Agent connects to.
The full model background lives in my Solar Mini 4 explainer, so I'll keep this part short.
🔥 Want the exact free Hermes brain setup I use?
Inside the AI Profit Boardroom I've got the full step-by-step video tutorials, the installable Agent OS, weekly coaching calls and 3,400+ members building real automations.
How to set up Solar Mini 4 in Hermes
There are four routes, and all of them end with a Hermes profile running the free model.
| Route | Steps | Best for |
|---|---|---|
| Hermes Cloud | Go to portal.nousresearch.com/cloud, create an agent, open the model section, scroll down, click Free and pick Upstage Solar Mini 4. | Running an agent with nothing installed locally. |
| Hermes Desktop | Create a new profile named after the model, refresh the models list, choose Nous Portal and pick Solar Mini 4 from the free models. | Testing a new brain without touching your main profile. |
| Hermes dashboard | Run hermes dashboard, open Models, click Set main model, type solar, choose Nous Portal and select the free variant. | Changing a profile's main model in under a minute. |
| Agent OS | Open the Hermes tab, pick the solar mini 4 profile and chat with Hermes directly. | Comparing several brains from one screen. |
The one thing you must do is pick the free variant.
Set main model shows a free Solar Mini 4 and a paid one next to each other, and the paid one bills you.
If the model is missing from your list, refresh it, or run hermes model --refresh to re-fetch every provider's live model list.
I keep one profile per model, and hermes profile create solar-mini-4 is how you make one from the terminal.
How Solar Mini 4 did in my first Hermes tests
I typed "test" into a fresh profile, and it replied fast and confirmed the session was live with memory switched off.
Then I asked for the latest AI automation news from the last seven days.
It came back quickly with a breakdown, sources and a "what stands out" section, including an Oracle launch I hadn't heard about.
That's a good result for a free model on a basic agentic task.
It's still not frontier level, and it isn't in the same league as Claude Opus 5.5 for hard work.
One viewer also told me it ran badly for them, so quality varies by task.
What GoldieBench measures, and what it doesn't
Every model on the board gets the same one-shot prompts to build games, pages, simulations and visuals.
Each build is rendered and scored from 0 to 10, and the scores you see are averages across those tasks.
That's a strong signal of how well a model plans and writes working code in one go.
It is not a direct test of agent loops like web research inside Hermes, so treat it as one input, not the whole answer.
Still, a brain that can't build reliably in one shot usually struggles on long multi-step agent jobs as well.
The best Hermes brains on the board right now
These are the top single models and set-ups on the live board, with the price each one lists.
| Model | GoldieBench avg | Scored tasks | Price as listed |
|---|---|---|---|
| Fusion (OpenRouter panel) | 8.59 | 47 | OpenRouter Fusion API pricing |
| Claude Opus 5 | 8.27 | 50 | $5 / $25 per M |
| Hermes MoA | 8.17 | 47 | Panel and aggregator calls via OpenRouter |
| GPT-5.6 Sol | 8.16 | 50 | $5 / $30 per M |
| MiniMax M3 | 7.97 | 47 | $0.30 / 1M input, $1.50 / 1M output |
| Kimi K3 | 7.89 | 50 | $3 / M in |
| GLM-5.2 | 7.77 | 47 | Open weights, free for individuals |
| Claude Opus 5.5 | 7.57 | 50 | $4 / $20 per M |
Hermes MoA is worth a special mention here, because it's Hermes Agent's own Mixture of Agents mode.
It scored 8.17 across 47 tasks, with 3 golds, 10 silvers and 2 bronzes.
That tells you Hermes itself can sit at the top of the board when you give it strong brains to combine.
For value, MiniMax M3 at 7.97 and GLM-5.2 at 7.77 are the ones I'd look at first.
The free Nous Portal models: which ones are benched?
When I recorded the video, the free list in Hermes Cloud showed six models next to Solar Mini 4.
Here's where each of them stands on GoldieBench today.
| Free model on Nous Portal | On GoldieBench? | Live avg | What that tells you |
|---|---|---|---|
| Upstage Solar Mini 4 | No | Not scored | It isn't benched yet, so all I can share is my own Hermes test above. |
| Meituan LongCat 2.0 | Yes, provisional | 8.12 across 4 tasks | It scored high on a small sample, so treat it as promising rather than proven. |
| Poolside Laguna XS 2.1 | Yes | 3.93 across 42 tasks | It struggles on one-shot builds, so keep it on light chores. |
| NVIDIA Nemotron | No | Not scored | It isn't benched, so test it on your own jobs. |
| inclusionAI Ling 3.1 Flash | No | Not scored | It isn't benched, so test it on your own jobs. |
| StepFun Step 3.7 Flash | No | Not scored | It isn't benched, so test it on your own jobs. |
The two free models I can give you numbers for sit at opposite ends of the board.
LongCat 2.0 averaged 8.12, but only 4 tasks are scored, which is why the board marks it provisional.
Laguna XS 2.1 averaged 3.93 across 42 tasks, which is a much bigger sample and a much weaker result.
Where Solar Mini 4 would fit
I haven't benched Solar Mini 4, so I'm not going to guess a number.
What I can do is tell you how I'd think about it, based on what the board shows for models of a similar shape.
It's a 3-billion-active mixture-of-experts model, and that size class is built for speed and price rather than depth.
On the local side of the board, Agents-A1 is also a 35B mixture-of-experts model with about 3B active, and it averages 4.83 across 45 tasks.
Agents-A1 runs locally and Solar Mini 4 is hosted, so they aren't the same thing, but it shows how hard one-shot builds are for small active-parameter models.
Solar Mini 4's own Intelligence Index score of 24 suggests it's strong for its size, and my research test backed that up.
So my working view is simple: put it in the fast lane for research, summaries and sorting, and keep the heavy builds on a model from the top table.
Free versus paid Hermes brains: what the gap really costs you
The board makes one thing very clear, which is that the top scores still belong to paid models.
Claude Opus 5 lists at $5 per million input tokens and $25 per million output tokens for its 8.27 average.
GPT-5.6 Sol lists at $5 per million input and $30 per million output for its 8.16 average.
MiniMax M3 lists at $0.30 per million input and $1.50 per million output, and it still averages 7.97.
That last row is the one I'd stare at if you're watching costs.
MiniMax M3 gets within about three tenths of a point of Claude Opus 5 at a small fraction of the input price.
A free model like Solar Mini 4 sits below all of that on price, because during the promo it costs nothing at all.
The trade is depth, and the board shows that trade over and over again in the smaller models.
So the smart move isn't picking one brain for everything.
The smart move is giving each Hermes profile the cheapest brain that can actually do its job.
How to test Solar Mini 4 in Hermes yourself
Until Solar Mini 4 has a GoldieBench score, your own test is the one that counts.
Here's the simple plan I use for any new free brain.
| Test | What to ask | What a pass looks like |
|---|---|---|
| Ping | Type test in a fresh profile. | It replies quickly and confirms the session is live. |
| Research | Ask for the latest news in your niche from the last 7 days with sources. | You get a structured brief with sources you can click and check. |
| Tool use | Ask it to read a folder and summarise the newest files. | It calls the file tool and summarises the right files. |
| Small build | Ask for a single-page HTML tool, such as a calculator. | The page opens and the main button works first time. |
| Follow-up | Ask it to change one thing in its last answer. | It keeps the context and makes only that change. |
The small build row is the closest thing to a mini GoldieBench task you can run at home.
If a model fails it, keep that model away from coding jobs in your Hermes setup.
If it passes all five, it has earned the fast lane.
Rate limits and switching between free brains
Free models are shared by a lot of people, so they get rate limited at busy times.
You'll notice slow replies or errors partway through a job when it happens.
Refresh the model list first, because a missing model is usually a stale cache rather than a removed model.
Then switch the profile to another free model, such as Step 3.7 Flash or Ling 3.1 Flash, and finish the job.
Switch back to Solar Mini 4 later, once the queue has calmed down.
This is why I keep one profile per model, because swapping brains becomes a single click.
The Hermes brain plan I'd use today
| Lane | Brain | Why |
|---|---|---|
| Fast and free | Solar Mini 4 (free variant), with Step 3.7 Flash or Ling 3.1 Flash as backups | It's free during the promo and quick on simple agentic jobs. |
| Free but heavier | LongCat 2.0 while it stays on the free list | It scored 8.12 on a small sample, so it's worth testing on bigger jobs. |
| Value | MiniMax M3 (7.97) or GLM-5.2 (7.77) | Near-frontier scores at a fraction of frontier prices. |
| Strongest | Claude Opus 5 (8.27) or Hermes MoA (8.17) | The best results on the board for hard planning and builds. |
| Local | Qwable 5 27B Coder (7.14) | The best local builder on the local board. |
Free models get rate limited, so the backups in the first lane matter.
If Solar Mini 4 slows down or vanishes, refresh the list, switch the profile, and come back later.
The free window is two weeks, so build each profile so the model is one setting you can change.
If you want my older, deeper ranking, the best LLMs for Hermes Agent post covers more of the board.
Also On Our Network
🌐 the step-by-step Solar Mini 4 Hermes setup for operators
🌐 how businesses use Solar Mini 4 inside Hermes
🌐 our scored verdict on Solar Mini 4 as a Hermes brain
🌐 the Solar Mini 4 Hermes quick start for your Agent OS
🌐 Julian's guide to running Hermes free forever
Pick the free variant, give it the fast lane, and watch the board for the day Solar Mini 4 Hermes gets a real GoldieBench score.
FAQ
Is Solar Mini 4 on the GoldieBench leaderboard?
No. Solar Mini 4 has not been benched on GoldieBench yet, so it has no score. Julian's own Hermes test showed a fast reply and a sourced seven-day news brief, but that is not a benchmark score.
How do I run Solar Mini 4 in Hermes for free?
Use Hermes Cloud at portal.nousresearch.com/cloud and pick Upstage Solar Mini 4 under Free, or run hermes dashboard, open Models, click Set main model, type solar, choose Nous Portal and select the free variant. It is free for two weeks on Nous Portal.
Which free Nous Portal models are benched on GoldieBench?
Meituan LongCat 2.0 is benched provisionally at 8.12 across 4 tasks, and Poolside Laguna XS 2.1 averages 3.93 across 42 tasks. Solar Mini 4, Nemotron, Ling 3.1 Flash and Step 3.7 Flash are not benched yet.
What is the best brain for Hermes Agent on GoldieBench?
Among single models, Claude Opus 5 leads at 8.27, with GPT-5.6 Sol at 8.16. Hermes MoA, Hermes Agent's Mixture of Agents mode, scores 8.17, and the OpenRouter Fusion panel tops the board at 8.59.
Is Solar Mini 4 good enough for serious Hermes work?
It is good for fast, basic agentic tasks like research and summaries, but it is not frontier level. Keep hard planning and builds on a stronger model such as Claude Opus 5, MiniMax M3 or GLM-5.2.
About Julian
I'm Julian Goldie — AI entrepreneur, SEO expert, and founder of the AI Profit Boardroom (3,400+ members).
I help business owners scale with AI agents, automation, and SEO.
400K+ YouTube subscribers. 7-figure AI agency (Goldie Agency). Daily training inside the Boardroom.
Related reading
→ The Best Local Model For Hermes Agent — Decided By 45 Real Builds
→ The Best LLMs For Hermes Agent, Ranked By Real Work
→ The Best Free AI Model For Hermes Agent (All Three $0 Lanes)
🌐 Sister-site take: read this on agentos.guide
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 3,400+ founders shipping with it every day all live inside the AI Profit Boardroom.
📺 Video notes + links to the tools 👉 AI Profit Boardroom
🆓 Get a FREE AI Course + Community + 1,000 AI Agents 👉 AI Money Lab
🎥 Learn how I make these videos 👉 agentos.guide