⭐ Get the Agent OS + join 3,400+ founders inside the AI Profit Boardroom → Join AIPB ($69/mo)

Solar Mini 4: The Specs, The Free Access, And How It Compares With Benchmarked Models

By Julian Goldie · 2026-10-07 · GoldieBench Blog

Solar Mini 4: The Specs, The Free Access, And How It Compares With Benchmarked Models — illustrated hero

Solar Mini 4 is Upstage's new compact mixture-of-experts model, and right now it's free for a two-week window inside Hermes Agent on Nous Portal.

I tested it on video, and it was fast, free and good at basic agentic jobs, but it isn't frontier level.

I also run GoldieBench, the leaderboard where 40 models have produced 1,609 one-shot demos across 50 tasks.

So this post does two things for you.

It explains exactly what Solar Mini 4 is and how to use it, and then it puts it in context with real scores from the board.

I'll be straight with you upfront: Solar Mini 4 has not been benched on GoldieBench yet, so I won't give it a score it hasn't earned.

What is Solar Mini 4?

Solar Mini 4 is a language model from Upstage, an AI lab based in South Korea.

Upstage describes it as a cost-efficient compact model for agentic use.

Here are the official specs, checked against Upstage's console docs, OpenRouter, Nous Portal and Artificial Analysis.

SpecOfficial figure
MakerUpstage, South Korea
ArchitectureMixture-of-experts with 35B total and 3B active parameters
Context window512K tokens on Upstage's docs, and 524,288 tokens on OpenRouter and Nous Portal
Max outputUp to 128K tokens on Upstage
LanguagesKorean, English and Japanese, input and output
CapabilitiesReasoning mode, tool calling, structured outputs and chat
Release22 September 2026, with a February 2026 training cut-off
WeightsProprietary, with no public download
Artificial Analysis Intelligence Index24

In my video I rounded the context to 500,000 tokens, and Upstage's official figure is 512K, so use that.

🔥 Want the exact Solar Mini 4 in Hermes setup I use?

Inside the AI Profit Boardroom I've got the full step-by-step video tutorials, the installable Agent OS, weekly coaching calls and 3,400+ members building real automations.

→ Get access here

Solar Mini 4 on third-party benchmarks

The main public score so far comes from Artificial Analysis.

It scores Solar Mini 4 at 24 on its Intelligence Index.

Their write-up says that's 6 points above Qwen 3.6 35B, which has the same 3B active parameters.

It's also 16 points above Upstage's previous flagship, Solar Pro 3, which scored 8.

In the video I framed it as scoring above models with ten times the active parameters, which is how the launch was pitched.

The same write-up has some honest weak spots that matter for agent work.

Artificial Analysis measurementSolar Mini 4 resultWhat it tells you
Intelligence Index24Strong for a 3B-active model, but well below the frontier
AA-LCR long-context reasoning83%Long documents are a genuine strength
Terminal-Bench 4.01%Terminal-heavy coding agents are a poor fit
AA-Omniscience knowledge18% accuracyDon't trust it on niche facts without sources
Output tokens per index taskAbout 88,000It's very verbose, which costs money once you pay per token

Those are Artificial Analysis numbers, not GoldieBench numbers, so read them as a separate test.

Is Solar Mini 4 on GoldieBench?

No, Solar Mini 4 is not on the GoldieBench leaderboard yet.

GoldieBench gives every model the same one-shot prompt to build a game, a page, a simulation or a visual, then scores the rendered result from 0 to 10.

Until Solar Mini 4 runs that same gauntlet, any score I gave it would be a guess, and I don't publish guesses.

What I can show you is how its closest neighbours have done on the board.

How small-active models score on GoldieBench

Solar Mini 4 runs about 3B active parameters, so the fairest comparisons are other models with a similar shape.

Two models on the board fit that description almost exactly.

ModelShapeGoldieBench averageTasks scored
Agents-A135B MoE with about 3B active4.8345
Laguna XS 2.133B MoE with 3B active3.9342

Laguna XS 2.1 is especially relevant, because it sits on the same Nous Portal free list as Solar Mini 4.

Neither one won a medal on the board, which tells you something useful about this class of model.

Small-active models are quick and cheap, but rich one-shot builds like 3D games still belong to bigger models.

That lines up with what I saw from Solar Mini 4 in Hermes, which was fast and useful on research but not frontier level.

You can see every local and lightweight score on the local models board.

Solar Mini 4 vs the other free models on Nous Portal

When I filtered Nous Portal by "Free", Solar Mini 4 sat next to NVIDIA Nemotron, Poolside Laguna XS 2.1, inclusionAI Ling 3.1 Flash and Ling 3.0 Flash, StepFun Step 3.7 Flash and Meituan LongCat 2.0.

Here's where each one stands on GoldieBench today.

Free model on Nous PortalOn GoldieBench?Live score
Upstage Solar Mini 4Not benched yetNo score
Meituan LongCat-2.0Yes, provisional8.12 across 4 tasks
Poolside Laguna XS 2.1Yes3.93 across 42 tasks
NVIDIA NemotronNot benched yetNo score
inclusionAI Ling 3.1 Flash and Ling 3.0 FlashNot benched yetNo score
StepFun Step 3.7 FlashNot benched yetNo score

LongCat-2.0's 8.12 is provisional because it has only been scored on 4 tasks, so treat it as an early signal rather than a final rank.

It's also a far bigger model, so it isn't a like-for-like comparison with a 3B-active model.

Solar Mini 4 vs the models people actually compare it with

In the video I said Solar Mini 4 isn't like Claude Opus 5.5, and the board shows why that comparison matters.

ModelGoldieBench averagePrice on the board
Claude Opus 58.27$5 / $25 per M
MiniMax M37.97$0.30 per M input, $1.50 per M output
Claude Opus 5.57.57$4 / $20 per M
Gemini 3.6 Flash7.08$1.50 per M input
Solar Mini 4Not benchedFree on Nous Portal for two weeks, or $0.10 in and $0.40 out per M on Upstage

The gap in price is huge, and so is the gap in build quality you should expect.

That's why I'd run Solar Mini 4 as a cheap worker lane, not as your main builder.

Why 3B active parameters matters for speed and quality

Total parameters tell you how much the model knows, and active parameters tell you how much of it works on each word.

Solar Mini 4 holds 35B parameters but only switches on about 3B for each step.

That's why it replies so quickly and why providers can sell it so cheaply.

The trade-off is depth, because a small active slice has less room for long chains of careful reasoning.

On GoldieBench, that trade-off shows up clearly in the scores of other small-active models.

Laguna XS 2.1 and Agents-A1 both run around 3B active, and both average under 5 out of 10 on one-shot builds.

Bigger models such as Claude Opus 5 and MiniMax M3 run far more compute per token and land near 8.

So when you pick Solar Mini 4, you're buying speed and price, and you're giving up some build quality.

For news checks, summaries and drafts, that's a trade worth making.

For a polished 3D game or a production web app, it isn't.

What Solar Mini 4 is good for, based on the numbers

The Artificial Analysis results and my own test point in the same direction.

Long-document work is its best lane, because it has a 512K window and scored 83% on AA-LCR.

Research sweeps are a good lane too, because it returned a sourced, well-organised news breakdown for me in Hermes.

Multilingual drafts are a natural fit, because Upstage supports Korean, English and Japanese.

Sorting and tagging jobs suit it, because it supports tool calling and structured outputs.

Terminal coding is its worst lane, because it scored just 1% on Terminal-Bench 4.0.

Niche factual questions are risky, because it scored 18% accuracy on AA-Omniscience, so always ask it for sources.

How to test Solar Mini 4 yourself, the GoldieBench way

You don't need my harness to run a fair test of your own.

First, pick three jobs you actually do every week, such as a news summary, a client email and a small web page.

Second, write one prompt for each job and keep it word for word the same across models.

Third, run the prompts on Solar Mini 4 and on the model you use today, each in its own Hermes profile.

Fourth, judge the outputs side by side without looking at which model wrote which.

Fifth, note the speed and the cost, because a slightly worse answer that's free and instant can still win.

That's the same idea behind GoldieBench, which is one prompt, many models and a blind, consistent judge.

How to use Solar Mini 4 in Hermes Agent

You can try it free in two ways inside Hermes.

The first way is Hermes Cloud, where you go to portal.nousresearch.com/cloud, create an agent, open the model section, click "Free" and pick Upstage Solar Mini 4.

The second way is Hermes Desktop, where you create a profile named "solar mini 4", choose Nous Portal, refresh the models list and pick it from the free models.

Then run hermes dashboard, go to Models, click "Set main model", type "solar", choose Nous Portal and pick the free variant.

This is the one step that matters most, because Nous Portal also lists a paid Solar Mini 4 at $0.05 per million input tokens and $0.20 per million output tokens.

The free model ID ends in ":free", and the paid one quietly bills you.

In the Agent OS, I open the Hermes tab and pick the solar mini 4 profile to chat with it directly.

The full step-by-step is in the Solar Mini 4 Hermes setup guide.

Upstage's own Playground is another free door, and Upstage says it needs no API key, but it's a browser chat rather than an agent.

What Solar Mini 4 did in my test

I typed "test" first, and it came back fast saying the session was live, with memory switched off.

Then I asked for AI automation news from the last seven days.

It came back quickly with a breakdown, sources and a "what stands out" section, and it flagged an Oracle launch I hadn't heard about.

That's a solid result for a free model on a basic agentic task.

One viewer said it ran badly for them, so quality clearly varies by task and prompt.

Limits to know before you rely on Solar Mini 4

Free models get rate limited, so keep another free model as a backup profile.

If it's missing from your list, refresh the models list before assuming it's gone.

The Nous Portal free window is two weeks, so don't build anything permanent on it being free.

The weights aren't released, so there's no local version for the local models board.

For a wider view of which brains suit Hermes, see the best LLMs for Hermes Agent.

Will Solar Mini 4 get a GoldieBench score?

I'd like to bench it, and the free window makes that easy to do.

When it runs, it will go through the same one-shot tasks, the same render checks and the same judge as every other model.

Until then, the honest answer is that its closest benched neighbours average between 3.93 and 4.83, and frontier models sit above 7.5.

Also On Our Network

🌐 the operator walkthrough for running Solar Mini 4 free

🌐 how businesses put Solar Mini 4 to work

🌐 a scored, feature-by-feature Solar Mini 4 verdict

🌐 the Solar Mini 4 quick start for an Agent OS

🌐 Julian's guide to running Hermes Agent free forever

Benched or not, the free window makes this the cheapest moment you'll ever get to test Solar Mini 4.

FAQ

What is Solar Mini 4?

Solar Mini 4 is a compact mixture-of-experts language model from Upstage in South Korea, with 35 billion total and 3 billion active parameters, a 512K context window and support for reasoning, tool calling and structured outputs.

Is Solar Mini 4 on GoldieBench?

Not yet. GoldieBench has not run Solar Mini 4 through its one-shot build tasks, so it has no GoldieBench score. Its closest benched neighbours are Agents-A1 at 4.83 and Laguna XS 2.1 at 3.93.

What benchmark scores does Solar Mini 4 have?

Artificial Analysis scores it 24 on its Intelligence Index, 83% on its AA-LCR long-context test and 1% on Terminal-Bench 4.0, and flags it as very verbose.

Is Solar Mini 4 free?

It is free for a two-week window on Nous Portal inside Hermes Agent, and Upstage's Playground lets you try it without an API key. Otherwise Upstage charges $0.10 per million input tokens and $0.40 per million output tokens.

How does Solar Mini 4 compare with Claude Opus 5.5?

Solar Mini 4 is much cheaper and faster, but it is not frontier level. Claude Opus 5.5 averages 7.57 on GoldieBench, while small 3B-active models on the board average between 3.93 and 4.83.

Which free Nous Portal models are on GoldieBench?

Laguna XS 2.1 scores 3.93 across 42 tasks and LongCat-2.0 has a provisional 8.12 across 4 tasks. Solar Mini 4, Nemotron, Ling Flash and Step Flash are not benched yet.

About Julian

I'm Julian Goldie — AI entrepreneur, SEO expert, and founder of the AI Profit Boardroom (3,400+ members).

I help business owners scale with AI agents, automation, and SEO.

400K+ YouTube subscribers. 7-figure AI agency (Goldie Agency). Daily training inside the Boardroom.

→ Get my best AI training inside the AI Profit Boardroom

Related reading

→ The Best Local Model For Hermes Agent — Decided By 45 Real Builds

→ The Best LLMs For Hermes Agent, Ranked By Real Work

→ The Best Free AI Model For Hermes Agent (All Three $0 Lanes)

🌐 Sister-site take: read this on agentos.guide

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 3,400+ founders shipping with it every day all live inside the AI Profit Boardroom.

3,400+founders
258documented wins
38countries
$69/momonthly

📺 Video notes + links to the tools 👉 AI Profit Boardroom

🆓 Get a FREE AI Course + Community + 1,000 AI Agents 👉 AI Money Lab

🎥 Learn how I make these videos 👉 agentos.guide

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 3,400+ founders shipping with it every day all live inside the AI Profit Boardroom.

3,400+founders
258documented wins
38countries
$69/momonthly