⭐ Get the Agent OS + join 3,400+ founders inside the AI Profit Boardroom → Join AIPB ($69/mo)

StarNet GitHub: Install The Repo, Then Pick A Brain With Real Benchmark Scores

By Julian Goldie · 2026-10-07 · GoldieBench Blog

StarNet GitHub: Install The Repo, Then Pick A Brain With Real Benchmark Scores — illustrated hero

The StarNet GitHub repo is github.com/androoAGI/starnet, and it holds the free, open-source code for StarNet, a pixel-art space station where your AI agents live and work.

StarNet does not ship with its own model.

You connect a brain, and that one choice decides how smart your agents are and how much they cost you.

I run GoldieBench, the leaderboard where 40 models have now produced 1,609 one-shot demos across 50 tasks.

So in this post I will show you what is in the repo and how to run it, then use real scores from the board to show you which brain to plug in.

What is in the StarNet GitHub repo?

StarNet describes itself as a local-first, gamified desktop harness for building and running real AI agent teams.

Every agent gets its own workspace, transcript, memory and permissions, and the station shows them working in real time.

These are the facts I checked in the repo on 7 October 2026.

ItemDetail
Repogithub.com/androoAGI/starnet, with signed installers on github.com/androoAGI/starnet-releases
AuthorAndrew Sims, who publishes as androoAGI
LicenceMIT for the code, with the StarNet name, logo and artwork reserved by the author
Versionv0.13.1, released on 4 October 2026
RequirementsNode.js 18 or newer and Git, plus Rust and Tauri only to build the desktop shell
PlatformsWindows and macOS installers, with Linux not offered as a public release
Default addresslocalhost:8787, or any port you set with the PORT setting
Clone sizeRoughly 1.8 GB

The repo had over 1,100 stars and more than 200 forks when I checked, and it was only created in June 2026.

🔥 Want the exact StarNet brain setup I use?

Inside the AI Profit Boardroom I've got the full step-by-step video tutorials, the installable Agent OS, weekly coaching calls and 3,400+ members building real automations.

→ Get access here

How to run StarNet from GitHub

These are the README commands, plus the one step my Mac needed.

git clone https://github.com/androoAGI/starnet.git
cd starnet
npm ci
node sidecar/index.js

Then open localhost:8787 in your browser and connect a brain on the first-run screen.

The README says the sidecar runs on Node core modules alone, but mine would not start properly until I ran npm ci.

I run mine with PORT=8788 node sidecar/index.js, because another app on my Mac already uses 8787.

To update, stop it, run git pull and npm ci, and start it again.

If you would rather skip the terminal, the Windows and Mac installers are on the StarNet releases page.

What GoldieBench can and cannot tell you about StarNet

I want to be straight about this before we look at numbers.

GoldieBench gives every model the same one-shot prompt to build a game, a page, a simulation or a visual, then scores the rendered result from 0 to 10.

StarNet asks a model to do something different, which is to act as an agent through many tool calls over a long session.

So a GoldieBench score is a measure of raw building ability, not a direct score for how well a model runs inside StarNet.

Nobody has benched StarNet itself on the board, and I am not going to pretend otherwise.

What the board does give you is a fair, same-prompt comparison of the brains StarNet lets you connect.

The StarNet brains, matched to real GoldieBench scores

StarNet's provider list is long, so I have matched the main sign-in options to the models they unlock on the board.

StarNet providerModel on GoldieBenchAvg scoreBuilds scoredNotes
Grok OAuth (SuperGrok or X Premium+)Grok 4.65.9520This is the model I picked for my own agents.
Grok OAuth (SuperGrok or X Premium+)Grok 4.77.1520It beat Grok 4.6 on 14 of the same 20 briefs.
Grok OAuth (SuperGrok or X Premium+)Grok (X Premium chat model)8.0943This is the X-native chat model rather than a numbered release.
Claude CLI (your Claude Code sign-in)Claude Opus 58.2750It is the highest single model on the board.
Claude CLI (your Claude Code sign-in)Claude Fable 58.1047It is strong, but it is the priciest Claude on the list.
Claude CLI (your Claude Code sign-in)Claude Opus 5.57.5750Every game it built was playtested before scoring.
Claude CLI (your Claude Code sign-in)Claude Sonnet 57.0147It is built for iterative, tool-using engineering work.
ChatGPT Codex sign-inGPT-5.6 Sol8.1650It sits just behind Opus 5.
Kimi OAuth (Kimi subscription)Kimi K37.8950It has a 1M context and is tuned for long agent runs.
Kimi OAuth (Kimi subscription)Kimi K2.77.4625It is the older Kimi on a flat plan.

Which Claude model you get through the Claude CLI depends on your Claude Code plan and settings, so check that before you compare.

Grok 4.6 or Grok 4.7 for StarNet?

In my StarNet video I picked Grok 4.6 over Grok 4.7.

That was a personal choice, because I prefer how Grok 4.6 talks.

The board tells a different story on building ability.

On the same 20 skill-infused game briefs, Grok 4.7 averaged 7.15 and Grok 4.6 averaged 5.95.

The Grok 4.7 vs Grok 4.6 head-to-head went to Grok 4.7 by 14 wins to 5, with one tie.

xAI lists both at $2 per million input tokens and $6 per million output on the standard API tier, but a Grok sign-in in StarNet runs on your subscription instead.

My advice is simple.

If your agents mostly chat, plan and research, pick the Grok that you like talking to.

If your agents build things, pick Grok 4.7, because the scores back it.

The best StarNet brain for each crew role

StarNet lets you recruit roles like strategist, chief of staff, researcher and marketer, and each agent can have its own model.

That means you do not need one brain for the whole station.

Crew roleWhat the role doesBrain I would try firstWhy
StrategistPlans, audits and big decisionsClaude Opus 5 (8.27) or GPT-5.6 Sol (8.16)These are the two highest single models on the board.
ResearcherLong reading and diggingKimi K3 (7.89)It has a 1M context window and is built for long agent runs.
MarketerCopy, hooks and ad variationsGrok 4.7 (7.15)It is fast to talk to and cheap on a Grok plan.
Chief of staffInbox, notes and daily briefClaude Sonnet 5 (7.01)It is tuned for tool-using work, which is most of what this role does.
Private jobsAnything sensitive that must stay offlineA local model through OllamaIt is free and private, but it is much weaker, as the next section shows.

Free and local brains for StarNet

StarNet can run with no key and no bill through Ollama.

The README suggests ollama pull qwen3:8b, then picking OLLAMA as the provider.

It also says to choose a model that supports tool calls, and it warns that an 8B model uses about 10 GB of graphics memory while it works.

Qwen3 8B is not on GoldieBench, so I cannot give you a score for that exact model.

Here is what the local models board says about the free options I have benched.

Local or free modelAvg scoreBuilds scoredHow it runs
Qwable 5 27B Coder7.1441It runs through MLX only, not Ollama, so StarNet's Ollama option cannot load it.
Agents-A14.8345It is an agent-tuned open MoE that runs free on a Mac.
Gemma 4 12B (MLX)3.9845It is the fast free engine on Ollama's MLX path.
Laguna XS 2.13.9342It has a free tier on OpenRouter, which StarNet also supports.
Qwythos 9B2.9842It is a small local model with a low build score.

StarNet does have a custom slot for any OpenAI-compatible endpoint, so a local MLX server might reach Qwable that way, but I have not tested it inside StarNet.

The gap is the real story.

The best free local builder scores 7.14, the next one down scores 4.83, and the cloud leaders sit above 8.

The README itself says local models are slower and rougher on long tasks.

If you want my full ranking of local brains for agents, read the best local model for Hermes Agent.

What each StarNet brain costs you

A sign-in brain and an API-key brain are billed in completely different ways.

When you sign in with Grok, Claude Code, Kimi or ChatGPT, StarNet runs on the subscription you already pay for, so there is no per-token bill inside StarNet.

When you paste an API key, every agent turn is billed per token, and a busy crew can burn through tokens quickly.

These are the list prices recorded on the model pages.

ModelList price per million tokensAvg score
Grok 4.6$2 input and $6 output on the standard tier5.95
Grok 4.7$2 input and $6 output on the standard tier7.15
Kimi K3$3 input on OpenRouter at launch7.89
Claude Sonnet 5$3 input and $15 output7.01
Claude Opus 5.5$4 input and $20 output7.57
Claude Opus 5$5 input and $25 output8.27
GPT-5.6 Sol$5 input and $30 output8.16

Grok 4.7 is the stand-out on value, because it scores 1.2 points higher than Grok 4.6 at exactly the same price.

StarNet also helps you keep the bill under control.

The README says spend, budgets and run history are saved on disk, and v0.13.1 made spending limits hold during retries as well as between turns.

Set a limit before you leave a crew running, whichever brain you pick.

What brain my own StarNet install runs on

My station runs on a hosted brain, not a local model.

The agents sign in with my Grok account and use Grok 4.6.

I do not run local models with StarNet, because I want the agents to be quick and sharp when I am showing them off.

That is the trade-off you are making.

A hosted brain gives you speed and quality, while a local brain gives you privacy and a zero bill.

A short tour of what the brain powers

This is the quick version, and my full walkthrough is the StarNet AI agent guide.

You create an agent with a name, an appearance, a personality such as warm, dry, composed or unhinged, and an approval level of ask for approval or full power.

You pick the model and a reasoning level of low, medium, high or max, and higher levels take longer.

The station view zooms and reframes, and everything you run lives in four docks.

Agents level up the more you use them, a bit like a Tamagotchi for an AI agent.

Each agent has Brief, Growth, History and Memories tabs, plus web and local-file permissions you can switch on or off.

The app catalogue adds Gmail, Notion, GitHub, flight search, physical letters, email and web search through Firecrawl.

The skill market and skill library add things like ad copy testing, calendar scheduling, a decision framework, making a plan, humanising content and handling a hard conversation.

The crew screen recruits roles, where Deploy rewrites an existing agent and Summon adds a new one beside it.

Tasks, quests, the work list, builds and the Systems menu all sit in the same station.

I noticed a slight delay when sending messages, so give each reply a moment.

My honest verdict on StarNet and its brains

I think StarNet is mostly fun to play with right now, and I do not see a massive number of practical use cases yet.

It is still the smoothest office-style agent harness I have tested, and it felt like a much more fun version of Paperclip.

If you already pay for Grok, Claude Code, Kimi or ChatGPT, plug that in first, because it costs you nothing extra.

If you are choosing from scratch, the board says Claude Opus 5 and GPT-5.6 Sol are the strongest brains, and Kimi K3 is the best value for long runs.

And if you live in Hermes, StarNet can import a Hermes or OpenClaw agent's persona, instructions, memory and model, though your API keys stay behind.

Also On Our Network

🌐 the step-by-step StarNet GitHub install for operators

🌐 seven ways a business can use the StarNet repo

🌐 my feature-by-feature StarNet verdict

🌐 running StarNet as a tab inside the Agent OS

🌐 how I wired Grok 4.7 into the Agent OS

Clone it, run npm ci, and connect the brain the board backs for the job, and you will get the most out of the StarNet GitHub repo.

FAQ

Where is the StarNet GitHub repo?

The source code is at github.com/androoAGI/starnet, published by Andrew Sims under the MIT License. Signed Windows and macOS installers are on github.com/androoAGI/starnet-releases. The latest version on 7 October 2026 was v0.13.1.

What is the best model to run StarNet on?

On GoldieBench, the strongest brains StarNet can connect are Claude Opus 5 (8.27) through the Claude CLI and GPT-5.6 Sol (8.16) through a ChatGPT sign-in. Kimi K3 (7.89) is a strong choice for long agent runs.

Is Grok 4.6 or Grok 4.7 better for StarNet?

On the same 20 game briefs, Grok 4.7 averaged 7.15 and Grok 4.6 averaged 5.95, and Grok 4.7 won the head-to-head 14 to 5. Julian still picked Grok 4.6 in his video because he prefers how it talks.

Can StarNet run on a free local model?

Yes. StarNet supports Ollama, and the README suggests qwen3:8b. Pick a model that supports tool calls. Benched local options score well below the cloud leaders, with Qwable 5 27B at 7.14 being MLX-only and Agents-A1 at 4.83.

Is StarNet itself on the GoldieBench leaderboard?

No. GoldieBench scores models on one-shot builds, and StarNet is a harness rather than a model. The scores show which brain to connect, not how StarNet performs as an app.

About Julian

I'm Julian Goldie — AI entrepreneur, SEO expert, and founder of the AI Profit Boardroom (3,400+ members).

I help business owners scale with AI agents, automation, and SEO.

400K+ YouTube subscribers. 7-figure AI agency (Goldie Agency). Daily training inside the Boardroom.

→ Get my best AI training inside the AI Profit Boardroom

Related reading

→ The Best Local Model For Hermes Agent — Decided By 45 Real Builds

→ The Best LLMs For Hermes Agent, Ranked By Real Work

→ The Best Free AI Model For Hermes Agent (All Three $0 Lanes)

🌐 Sister-site take: read this on agentos.guide

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 3,400+ founders shipping with it every day all live inside the AI Profit Boardroom.

3,400+founders
258documented wins
38countries
$69/momonthly

📺 Video notes + links to the tools 👉 AI Profit Boardroom

🆓 Get a FREE AI Course + Community + 1,000 AI Agents 👉 AI Money Lab

🎥 Learn how I make these videos 👉 agentos.guide

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 3,400+ founders shipping with it every day all live inside the AI Profit Boardroom.

3,400+founders
258documented wins
38countries
$69/momonthly