Solar Mini 4: The Specs, The Free Access, And How It Compares With Benchmarked Models
By Julian Goldie · 2026-10-07 · GoldieBench Blog
Solar Mini 4 is Upstage's new compact mixture-of-experts model, and right now it's free for a two-week window inside Hermes Agent on Nous Portal.
I tested it on video, and it was fast, free and good at basic agentic jobs, but it isn't frontier level.
I also run GoldieBench, the leaderboard where 40 models have produced 1,609 one-shot demos across 50 tasks.
So this post does two things for you.
It explains exactly what Solar Mini 4 is and how to use it, and then it puts it in context with real scores from the board.
I'll be straight with you upfront: Solar Mini 4 has not been benched on GoldieBench yet, so I won't give it a score it hasn't earned.
What is Solar Mini 4?
Solar Mini 4 is a language model from Upstage, an AI lab based in South Korea.
Upstage describes it as a cost-efficient compact model for agentic use.
Here are the official specs, checked against Upstage's console docs, OpenRouter, Nous Portal and Artificial Analysis.
| Spec | Official figure |
|---|---|
| Maker | Upstage, South Korea |
| Architecture | Mixture-of-experts with 35B total and 3B active parameters |
| Context window | 512K tokens on Upstage's docs, and 524,288 tokens on OpenRouter and Nous Portal |
| Max output | Up to 128K tokens on Upstage |
| Languages | Korean, English and Japanese, input and output |
| Capabilities | Reasoning mode, tool calling, structured outputs and chat |
| Release | 22 September 2026, with a February 2026 training cut-off |
| Weights | Proprietary, with no public download |
| Artificial Analysis Intelligence Index | 24 |
In my video I rounded the context to 500,000 tokens, and Upstage's official figure is 512K, so use that.
🔥 Want the exact Solar Mini 4 in Hermes setup I use?
Inside the AI Profit Boardroom I've got the full step-by-step video tutorials, the installable Agent OS, weekly coaching calls and 3,400+ members building real automations.
Solar Mini 4 on third-party benchmarks
The main public score so far comes from Artificial Analysis.
It scores Solar Mini 4 at 24 on its Intelligence Index.
Their write-up says that's 6 points above Qwen 3.6 35B, which has the same 3B active parameters.
It's also 16 points above Upstage's previous flagship, Solar Pro 3, which scored 8.
In the video I framed it as scoring above models with ten times the active parameters, which is how the launch was pitched.
The same write-up has some honest weak spots that matter for agent work.
| Artificial Analysis measurement | Solar Mini 4 result | What it tells you |
|---|---|---|
| Intelligence Index | 24 | Strong for a 3B-active model, but well below the frontier |
| AA-LCR long-context reasoning | 83% | Long documents are a genuine strength |
| Terminal-Bench 4.0 | 1% | Terminal-heavy coding agents are a poor fit |
| AA-Omniscience knowledge | 18% accuracy | Don't trust it on niche facts without sources |
| Output tokens per index task | About 88,000 | It's very verbose, which costs money once you pay per token |
Those are Artificial Analysis numbers, not GoldieBench numbers, so read them as a separate test.
Is Solar Mini 4 on GoldieBench?
No, Solar Mini 4 is not on the GoldieBench leaderboard yet.
GoldieBench gives every model the same one-shot prompt to build a game, a page, a simulation or a visual, then scores the rendered result from 0 to 10.
Until Solar Mini 4 runs that same gauntlet, any score I gave it would be a guess, and I don't publish guesses.
What I can show you is how its closest neighbours have done on the board.
How small-active models score on GoldieBench
Solar Mini 4 runs about 3B active parameters, so the fairest comparisons are other models with a similar shape.
Two models on the board fit that description almost exactly.
| Model | Shape | GoldieBench average | Tasks scored |
|---|---|---|---|
| Agents-A1 | 35B MoE with about 3B active | 4.83 | 45 |
| Laguna XS 2.1 | 33B MoE with 3B active | 3.93 | 42 |
Laguna XS 2.1 is especially relevant, because it sits on the same Nous Portal free list as Solar Mini 4.
Neither one won a medal on the board, which tells you something useful about this class of model.
Small-active models are quick and cheap, but rich one-shot builds like 3D games still belong to bigger models.
That lines up with what I saw from Solar Mini 4 in Hermes, which was fast and useful on research but not frontier level.
You can see every local and lightweight score on the local models board.
Solar Mini 4 vs the other free models on Nous Portal
When I filtered Nous Portal by "Free", Solar Mini 4 sat next to NVIDIA Nemotron, Poolside Laguna XS 2.1, inclusionAI Ling 3.1 Flash and Ling 3.0 Flash, StepFun Step 3.7 Flash and Meituan LongCat 2.0.
Here's where each one stands on GoldieBench today.
| Free model on Nous Portal | On GoldieBench? | Live score |
|---|---|---|
| Upstage Solar Mini 4 | Not benched yet | No score |
| Meituan LongCat-2.0 | Yes, provisional | 8.12 across 4 tasks |
| Poolside Laguna XS 2.1 | Yes | 3.93 across 42 tasks |
| NVIDIA Nemotron | Not benched yet | No score |
| inclusionAI Ling 3.1 Flash and Ling 3.0 Flash | Not benched yet | No score |
| StepFun Step 3.7 Flash | Not benched yet | No score |
LongCat-2.0's 8.12 is provisional because it has only been scored on 4 tasks, so treat it as an early signal rather than a final rank.
It's also a far bigger model, so it isn't a like-for-like comparison with a 3B-active model.
Solar Mini 4 vs the models people actually compare it with
In the video I said Solar Mini 4 isn't like Claude Opus 5.5, and the board shows why that comparison matters.
| Model | GoldieBench average | Price on the board |
|---|---|---|
| Claude Opus 5 | 8.27 | $5 / $25 per M |
| MiniMax M3 | 7.97 | $0.30 per M input, $1.50 per M output |
| Claude Opus 5.5 | 7.57 | $4 / $20 per M |
| Gemini 3.6 Flash | 7.08 | $1.50 per M input |
| Solar Mini 4 | Not benched | Free on Nous Portal for two weeks, or $0.10 in and $0.40 out per M on Upstage |
The gap in price is huge, and so is the gap in build quality you should expect.
That's why I'd run Solar Mini 4 as a cheap worker lane, not as your main builder.
Why 3B active parameters matters for speed and quality
Total parameters tell you how much the model knows, and active parameters tell you how much of it works on each word.
Solar Mini 4 holds 35B parameters but only switches on about 3B for each step.
That's why it replies so quickly and why providers can sell it so cheaply.
The trade-off is depth, because a small active slice has less room for long chains of careful reasoning.
On GoldieBench, that trade-off shows up clearly in the scores of other small-active models.
Laguna XS 2.1 and Agents-A1 both run around 3B active, and both average under 5 out of 10 on one-shot builds.
Bigger models such as Claude Opus 5 and MiniMax M3 run far more compute per token and land near 8.
So when you pick Solar Mini 4, you're buying speed and price, and you're giving up some build quality.
For news checks, summaries and drafts, that's a trade worth making.
For a polished 3D game or a production web app, it isn't.
What Solar Mini 4 is good for, based on the numbers
The Artificial Analysis results and my own test point in the same direction.
Long-document work is its best lane, because it has a 512K window and scored 83% on AA-LCR.
Research sweeps are a good lane too, because it returned a sourced, well-organised news breakdown for me in Hermes.
Multilingual drafts are a natural fit, because Upstage supports Korean, English and Japanese.
Sorting and tagging jobs suit it, because it supports tool calling and structured outputs.
Terminal coding is its worst lane, because it scored just 1% on Terminal-Bench 4.0.
Niche factual questions are risky, because it scored 18% accuracy on AA-Omniscience, so always ask it for sources.
How to test Solar Mini 4 yourself, the GoldieBench way
You don't need my harness to run a fair test of your own.
First, pick three jobs you actually do every week, such as a news summary, a client email and a small web page.
Second, write one prompt for each job and keep it word for word the same across models.
Third, run the prompts on Solar Mini 4 and on the model you use today, each in its own Hermes profile.
Fourth, judge the outputs side by side without looking at which model wrote which.
Fifth, note the speed and the cost, because a slightly worse answer that's free and instant can still win.
That's the same idea behind GoldieBench, which is one prompt, many models and a blind, consistent judge.
How to use Solar Mini 4 in Hermes Agent
You can try it free in two ways inside Hermes.
The first way is Hermes Cloud, where you go to portal.nousresearch.com/cloud, create an agent, open the model section, click "Free" and pick Upstage Solar Mini 4.
The second way is Hermes Desktop, where you create a profile named "solar mini 4", choose Nous Portal, refresh the models list and pick it from the free models.
Then run hermes dashboard, go to Models, click "Set main model", type "solar", choose Nous Portal and pick the free variant.
This is the one step that matters most, because Nous Portal also lists a paid Solar Mini 4 at $0.05 per million input tokens and $0.20 per million output tokens.
The free model ID ends in ":free", and the paid one quietly bills you.
In the Agent OS, I open the Hermes tab and pick the solar mini 4 profile to chat with it directly.
The full step-by-step is in the Solar Mini 4 Hermes setup guide.
Upstage's own Playground is another free door, and Upstage says it needs no API key, but it's a browser chat rather than an agent.
What Solar Mini 4 did in my test
I typed "test" first, and it came back fast saying the session was live, with memory switched off.
Then I asked for AI automation news from the last seven days.
It came back quickly with a breakdown, sources and a "what stands out" section, and it flagged an Oracle launch I hadn't heard about.
That's a solid result for a free model on a basic agentic task.
One viewer said it ran badly for them, so quality clearly varies by task and prompt.
Limits to know before you rely on Solar Mini 4
Free models get rate limited, so keep another free model as a backup profile.
If it's missing from your list, refresh the models list before assuming it's gone.
The Nous Portal free window is two weeks, so don't build anything permanent on it being free.
The weights aren't released, so there's no local version for the local models board.
For a wider view of which brains suit Hermes, see the best LLMs for Hermes Agent.
Will Solar Mini 4 get a GoldieBench score?
I'd like to bench it, and the free window makes that easy to do.
When it runs, it will go through the same one-shot tasks, the same render checks and the same judge as every other model.
Until then, the honest answer is that its closest benched neighbours average between 3.93 and 4.83, and frontier models sit above 7.5.
Also On Our Network
🌐 the operator walkthrough for running Solar Mini 4 free
🌐 how businesses put Solar Mini 4 to work
🌐 a scored, feature-by-feature Solar Mini 4 verdict
🌐 the Solar Mini 4 quick start for an Agent OS
🌐 Julian's guide to running Hermes Agent free forever
Benched or not, the free window makes this the cheapest moment you'll ever get to test Solar Mini 4.
FAQ
What is Solar Mini 4?
Solar Mini 4 is a compact mixture-of-experts language model from Upstage in South Korea, with 35 billion total and 3 billion active parameters, a 512K context window and support for reasoning, tool calling and structured outputs.
Is Solar Mini 4 on GoldieBench?
Not yet. GoldieBench has not run Solar Mini 4 through its one-shot build tasks, so it has no GoldieBench score. Its closest benched neighbours are Agents-A1 at 4.83 and Laguna XS 2.1 at 3.93.
What benchmark scores does Solar Mini 4 have?
Artificial Analysis scores it 24 on its Intelligence Index, 83% on its AA-LCR long-context test and 1% on Terminal-Bench 4.0, and flags it as very verbose.
Is Solar Mini 4 free?
It is free for a two-week window on Nous Portal inside Hermes Agent, and Upstage's Playground lets you try it without an API key. Otherwise Upstage charges $0.10 per million input tokens and $0.40 per million output tokens.
How does Solar Mini 4 compare with Claude Opus 5.5?
Solar Mini 4 is much cheaper and faster, but it is not frontier level. Claude Opus 5.5 averages 7.57 on GoldieBench, while small 3B-active models on the board average between 3.93 and 4.83.
Which free Nous Portal models are on GoldieBench?
Laguna XS 2.1 scores 3.93 across 42 tasks and LongCat-2.0 has a provisional 8.12 across 4 tasks. Solar Mini 4, Nemotron, Ling Flash and Step Flash are not benched yet.
About Julian
I'm Julian Goldie — AI entrepreneur, SEO expert, and founder of the AI Profit Boardroom (3,400+ members).
I help business owners scale with AI agents, automation, and SEO.
400K+ YouTube subscribers. 7-figure AI agency (Goldie Agency). Daily training inside the Boardroom.
Related reading
→ The Best Local Model For Hermes Agent — Decided By 45 Real Builds
→ The Best LLMs For Hermes Agent, Ranked By Real Work
→ The Best Free AI Model For Hermes Agent (All Three $0 Lanes)
🌐 Sister-site take: read this on agentos.guide
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 3,400+ founders shipping with it every day all live inside the AI Profit Boardroom.
📺 Video notes + links to the tools 👉 AI Profit Boardroom
🆓 Get a FREE AI Course + Community + 1,000 AI Agents 👉 AI Money Lab
🎥 Learn how I make these videos 👉 agentos.guide