StarNet GitHub: Install The Repo, Then Pick A Brain With Real Benchmark Scores
By Julian Goldie · 2026-10-07 · GoldieBench Blog
The StarNet GitHub repo is github.com/androoAGI/starnet, and it holds the free, open-source code for StarNet, a pixel-art space station where your AI agents live and work.
StarNet does not ship with its own model.
You connect a brain, and that one choice decides how smart your agents are and how much they cost you.
I run GoldieBench, the leaderboard where 40 models have now produced 1,609 one-shot demos across 50 tasks.
So in this post I will show you what is in the repo and how to run it, then use real scores from the board to show you which brain to plug in.
What is in the StarNet GitHub repo?
StarNet describes itself as a local-first, gamified desktop harness for building and running real AI agent teams.
Every agent gets its own workspace, transcript, memory and permissions, and the station shows them working in real time.
These are the facts I checked in the repo on 7 October 2026.
| Item | Detail |
|---|---|
| Repo | github.com/androoAGI/starnet, with signed installers on github.com/androoAGI/starnet-releases |
| Author | Andrew Sims, who publishes as androoAGI |
| Licence | MIT for the code, with the StarNet name, logo and artwork reserved by the author |
| Version | v0.13.1, released on 4 October 2026 |
| Requirements | Node.js 18 or newer and Git, plus Rust and Tauri only to build the desktop shell |
| Platforms | Windows and macOS installers, with Linux not offered as a public release |
| Default address | localhost:8787, or any port you set with the PORT setting |
| Clone size | Roughly 1.8 GB |
The repo had over 1,100 stars and more than 200 forks when I checked, and it was only created in June 2026.
🔥 Want the exact StarNet brain setup I use?
Inside the AI Profit Boardroom I've got the full step-by-step video tutorials, the installable Agent OS, weekly coaching calls and 3,400+ members building real automations.
How to run StarNet from GitHub
These are the README commands, plus the one step my Mac needed.
git clone https://github.com/androoAGI/starnet.git
cd starnet
npm ci
node sidecar/index.js
Then open localhost:8787 in your browser and connect a brain on the first-run screen.
The README says the sidecar runs on Node core modules alone, but mine would not start properly until I ran npm ci.
I run mine with PORT=8788 node sidecar/index.js, because another app on my Mac already uses 8787.
To update, stop it, run git pull and npm ci, and start it again.
If you would rather skip the terminal, the Windows and Mac installers are on the StarNet releases page.
What GoldieBench can and cannot tell you about StarNet
I want to be straight about this before we look at numbers.
GoldieBench gives every model the same one-shot prompt to build a game, a page, a simulation or a visual, then scores the rendered result from 0 to 10.
StarNet asks a model to do something different, which is to act as an agent through many tool calls over a long session.
So a GoldieBench score is a measure of raw building ability, not a direct score for how well a model runs inside StarNet.
Nobody has benched StarNet itself on the board, and I am not going to pretend otherwise.
What the board does give you is a fair, same-prompt comparison of the brains StarNet lets you connect.
The StarNet brains, matched to real GoldieBench scores
StarNet's provider list is long, so I have matched the main sign-in options to the models they unlock on the board.
| StarNet provider | Model on GoldieBench | Avg score | Builds scored | Notes |
|---|---|---|---|---|
| Grok OAuth (SuperGrok or X Premium+) | Grok 4.6 | 5.95 | 20 | This is the model I picked for my own agents. |
| Grok OAuth (SuperGrok or X Premium+) | Grok 4.7 | 7.15 | 20 | It beat Grok 4.6 on 14 of the same 20 briefs. |
| Grok OAuth (SuperGrok or X Premium+) | Grok (X Premium chat model) | 8.09 | 43 | This is the X-native chat model rather than a numbered release. |
| Claude CLI (your Claude Code sign-in) | Claude Opus 5 | 8.27 | 50 | It is the highest single model on the board. |
| Claude CLI (your Claude Code sign-in) | Claude Fable 5 | 8.10 | 47 | It is strong, but it is the priciest Claude on the list. |
| Claude CLI (your Claude Code sign-in) | Claude Opus 5.5 | 7.57 | 50 | Every game it built was playtested before scoring. |
| Claude CLI (your Claude Code sign-in) | Claude Sonnet 5 | 7.01 | 47 | It is built for iterative, tool-using engineering work. |
| ChatGPT Codex sign-in | GPT-5.6 Sol | 8.16 | 50 | It sits just behind Opus 5. |
| Kimi OAuth (Kimi subscription) | Kimi K3 | 7.89 | 50 | It has a 1M context and is tuned for long agent runs. |
| Kimi OAuth (Kimi subscription) | Kimi K2.7 | 7.46 | 25 | It is the older Kimi on a flat plan. |
Which Claude model you get through the Claude CLI depends on your Claude Code plan and settings, so check that before you compare.
Grok 4.6 or Grok 4.7 for StarNet?
In my StarNet video I picked Grok 4.6 over Grok 4.7.
That was a personal choice, because I prefer how Grok 4.6 talks.
The board tells a different story on building ability.
On the same 20 skill-infused game briefs, Grok 4.7 averaged 7.15 and Grok 4.6 averaged 5.95.
The Grok 4.7 vs Grok 4.6 head-to-head went to Grok 4.7 by 14 wins to 5, with one tie.
xAI lists both at $2 per million input tokens and $6 per million output on the standard API tier, but a Grok sign-in in StarNet runs on your subscription instead.
My advice is simple.
If your agents mostly chat, plan and research, pick the Grok that you like talking to.
If your agents build things, pick Grok 4.7, because the scores back it.
The best StarNet brain for each crew role
StarNet lets you recruit roles like strategist, chief of staff, researcher and marketer, and each agent can have its own model.
That means you do not need one brain for the whole station.
| Crew role | What the role does | Brain I would try first | Why |
|---|---|---|---|
| Strategist | Plans, audits and big decisions | Claude Opus 5 (8.27) or GPT-5.6 Sol (8.16) | These are the two highest single models on the board. |
| Researcher | Long reading and digging | Kimi K3 (7.89) | It has a 1M context window and is built for long agent runs. |
| Marketer | Copy, hooks and ad variations | Grok 4.7 (7.15) | It is fast to talk to and cheap on a Grok plan. |
| Chief of staff | Inbox, notes and daily brief | Claude Sonnet 5 (7.01) | It is tuned for tool-using work, which is most of what this role does. |
| Private jobs | Anything sensitive that must stay offline | A local model through Ollama | It is free and private, but it is much weaker, as the next section shows. |
Free and local brains for StarNet
StarNet can run with no key and no bill through Ollama.
The README suggests ollama pull qwen3:8b, then picking OLLAMA as the provider.
It also says to choose a model that supports tool calls, and it warns that an 8B model uses about 10 GB of graphics memory while it works.
Qwen3 8B is not on GoldieBench, so I cannot give you a score for that exact model.
Here is what the local models board says about the free options I have benched.
| Local or free model | Avg score | Builds scored | How it runs |
|---|---|---|---|
| Qwable 5 27B Coder | 7.14 | 41 | It runs through MLX only, not Ollama, so StarNet's Ollama option cannot load it. |
| Agents-A1 | 4.83 | 45 | It is an agent-tuned open MoE that runs free on a Mac. |
| Gemma 4 12B (MLX) | 3.98 | 45 | It is the fast free engine on Ollama's MLX path. |
| Laguna XS 2.1 | 3.93 | 42 | It has a free tier on OpenRouter, which StarNet also supports. |
| Qwythos 9B | 2.98 | 42 | It is a small local model with a low build score. |
StarNet does have a custom slot for any OpenAI-compatible endpoint, so a local MLX server might reach Qwable that way, but I have not tested it inside StarNet.
The gap is the real story.
The best free local builder scores 7.14, the next one down scores 4.83, and the cloud leaders sit above 8.
The README itself says local models are slower and rougher on long tasks.
If you want my full ranking of local brains for agents, read the best local model for Hermes Agent.
What each StarNet brain costs you
A sign-in brain and an API-key brain are billed in completely different ways.
When you sign in with Grok, Claude Code, Kimi or ChatGPT, StarNet runs on the subscription you already pay for, so there is no per-token bill inside StarNet.
When you paste an API key, every agent turn is billed per token, and a busy crew can burn through tokens quickly.
These are the list prices recorded on the model pages.
| Model | List price per million tokens | Avg score |
|---|---|---|
| Grok 4.6 | $2 input and $6 output on the standard tier | 5.95 |
| Grok 4.7 | $2 input and $6 output on the standard tier | 7.15 |
| Kimi K3 | $3 input on OpenRouter at launch | 7.89 |
| Claude Sonnet 5 | $3 input and $15 output | 7.01 |
| Claude Opus 5.5 | $4 input and $20 output | 7.57 |
| Claude Opus 5 | $5 input and $25 output | 8.27 |
| GPT-5.6 Sol | $5 input and $30 output | 8.16 |
Grok 4.7 is the stand-out on value, because it scores 1.2 points higher than Grok 4.6 at exactly the same price.
StarNet also helps you keep the bill under control.
The README says spend, budgets and run history are saved on disk, and v0.13.1 made spending limits hold during retries as well as between turns.
Set a limit before you leave a crew running, whichever brain you pick.
What brain my own StarNet install runs on
My station runs on a hosted brain, not a local model.
The agents sign in with my Grok account and use Grok 4.6.
I do not run local models with StarNet, because I want the agents to be quick and sharp when I am showing them off.
That is the trade-off you are making.
A hosted brain gives you speed and quality, while a local brain gives you privacy and a zero bill.
A short tour of what the brain powers
This is the quick version, and my full walkthrough is the StarNet AI agent guide.
You create an agent with a name, an appearance, a personality such as warm, dry, composed or unhinged, and an approval level of ask for approval or full power.
You pick the model and a reasoning level of low, medium, high or max, and higher levels take longer.
The station view zooms and reframes, and everything you run lives in four docks.
Agents level up the more you use them, a bit like a Tamagotchi for an AI agent.
Each agent has Brief, Growth, History and Memories tabs, plus web and local-file permissions you can switch on or off.
The app catalogue adds Gmail, Notion, GitHub, flight search, physical letters, email and web search through Firecrawl.
The skill market and skill library add things like ad copy testing, calendar scheduling, a decision framework, making a plan, humanising content and handling a hard conversation.
The crew screen recruits roles, where Deploy rewrites an existing agent and Summon adds a new one beside it.
Tasks, quests, the work list, builds and the Systems menu all sit in the same station.
I noticed a slight delay when sending messages, so give each reply a moment.
My honest verdict on StarNet and its brains
I think StarNet is mostly fun to play with right now, and I do not see a massive number of practical use cases yet.
It is still the smoothest office-style agent harness I have tested, and it felt like a much more fun version of Paperclip.
If you already pay for Grok, Claude Code, Kimi or ChatGPT, plug that in first, because it costs you nothing extra.
If you are choosing from scratch, the board says Claude Opus 5 and GPT-5.6 Sol are the strongest brains, and Kimi K3 is the best value for long runs.
And if you live in Hermes, StarNet can import a Hermes or OpenClaw agent's persona, instructions, memory and model, though your API keys stay behind.
Also On Our Network
🌐 the step-by-step StarNet GitHub install for operators
🌐 seven ways a business can use the StarNet repo
🌐 my feature-by-feature StarNet verdict
🌐 running StarNet as a tab inside the Agent OS
🌐 how I wired Grok 4.7 into the Agent OS
Clone it, run npm ci, and connect the brain the board backs for the job, and you will get the most out of the StarNet GitHub repo.
FAQ
Where is the StarNet GitHub repo?
The source code is at github.com/androoAGI/starnet, published by Andrew Sims under the MIT License. Signed Windows and macOS installers are on github.com/androoAGI/starnet-releases. The latest version on 7 October 2026 was v0.13.1.
What is the best model to run StarNet on?
On GoldieBench, the strongest brains StarNet can connect are Claude Opus 5 (8.27) through the Claude CLI and GPT-5.6 Sol (8.16) through a ChatGPT sign-in. Kimi K3 (7.89) is a strong choice for long agent runs.
Is Grok 4.6 or Grok 4.7 better for StarNet?
On the same 20 game briefs, Grok 4.7 averaged 7.15 and Grok 4.6 averaged 5.95, and Grok 4.7 won the head-to-head 14 to 5. Julian still picked Grok 4.6 in his video because he prefers how it talks.
Can StarNet run on a free local model?
Yes. StarNet supports Ollama, and the README suggests qwen3:8b. Pick a model that supports tool calls. Benched local options score well below the cloud leaders, with Qwable 5 27B at 7.14 being MLX-only and Agents-A1 at 4.83.
Is StarNet itself on the GoldieBench leaderboard?
No. GoldieBench scores models on one-shot builds, and StarNet is a harness rather than a model. The scores show which brain to connect, not how StarNet performs as an app.
About Julian
I'm Julian Goldie — AI entrepreneur, SEO expert, and founder of the AI Profit Boardroom (3,400+ members).
I help business owners scale with AI agents, automation, and SEO.
400K+ YouTube subscribers. 7-figure AI agency (Goldie Agency). Daily training inside the Boardroom.
Related reading
→ The Best Local Model For Hermes Agent — Decided By 45 Real Builds
→ The Best LLMs For Hermes Agent, Ranked By Real Work
→ The Best Free AI Model For Hermes Agent (All Three $0 Lanes)
🌐 Sister-site take: read this on agentos.guide
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 3,400+ founders shipping with it every day all live inside the AI Profit Boardroom.
📺 Video notes + links to the tools 👉 AI Profit Boardroom
🆓 Get a FREE AI Course + Community + 1,000 AI Agents 👉 AI Money Lab
🎥 Learn how I make these videos 👉 agentos.guide