⭐ Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)

Can You Run Jev Locally? The Honest Answer, Plus The Local Stack The Bench Supports

By Julian Goldie · 2026-10-05 · GoldieBench Blog

Can You Run Jev Locally? The Honest Answer, Plus The Local Stack The Bench Supports — illustrated hero

No, you cannot run Jev locally, because TypeSafe has not released Jev's model weights and it only runs as a hosted service.

What you can run locally is a Jev-shaped decision model, plus a local builder model underneath it.

I benchmark local models on GoldieBench every week, so this post is about the second half that everyone skips.

A local decision layer is pointless if the model doing the actual work is weak.

Why you cannot run Jev locally today

Jev exists only behind TypeSafe's API and the gateways that resell it, such as OpenCode Zen, Vercel AI Gateway and OpenRouter.

The SDKs are open source under the MIT licence, but an SDK is only the code that talks to the hosted model.

So any page offering a Jev download is either selling something else or confusing the SDK with the model.

That saves you an afternoon of hunting.

🔥 Want the exact local decision-model setup I use?

Inside the AI Profit Boardroom I've got the full step-by-step video tutorials, the installable Agent OS, weekly coaching calls and 4,000+ members building real automations.

→ Get access here

What runs locally instead

Here is the honest list of Jev-style decision models you can put on your own machine.

OptionRuns locally?What I know
TypeSafe JevNo, hosted onlyGot 54 of 60 right on my email test
Laya (pip install laya)Yes, I ran 0.3.4 on my MacAbout 20 ms per English question, but untrained it scored 28 and 36 of 60 on the same emails
Cloudflare ClefYes, weights on Hugging FaceReleased 1 October 2026, Apache 2.0, built on Qwen 3.8-27B, and Cloudflare says it is Jev-API compatible
Cloudflare Clef-flashYes, weights on Hugging FaceA 9B sibling built on Qwen 3.5-9B, a far easier fit on a laptop

I have installed and tested Laya myself.

I have not run Clef on my own machine yet, so I am not going to give you a speed number for it.

The accuracy gap you need to know about

Fast is not the same as right.

TestHosted JevLaya, untrained
60 real-style emails54 of 6028 (English model) and 36 (typed-decisions model)
24 contact-form leads24 of 2416 of 24
Pick 1 of 5 websites for a brief10 of 108 of 10
77-category test0.8700.425

The good news is that Laya did not fake confidence when it was guessing.

None of its 60 email answers came back at 0.85 confidence or higher.

That means a confidence threshold still works with a local model, even when the answers are weaker.

Run Jev locally? Pair the decision model with a local builder

This is where GoldieBench comes in.

A local decision model only decides where work goes.

Something else has to do the work, and on a fully offline stack that something is a local model.

Here is how the local builders score on the same 50-task gauntlet the frontier models run.

Local modelGoldieBench avgNotes from the board
Qwable 5 27B Coder7.14Best local builder, MLX only, built on Qwen3.6-27B
Agents-A14.83Agent-tuned open MoE built for long tool work
Gemma 4 12B Coder4.25Free offline coder
Gemma 4 12B MLX3.98The fast lightweight engine
Qwythos 9B2.98A local writer with a 1M-token context

For context, the top of the board is Claude Opus 5 at 8.27 and GPT-5.6 Sol at 8.16.

So the best local builder gets you most of the way, and the rest drop off fast.

One honest caveat about Clef: the Qwen 3.8 on our board scores 8.10, but that is Alibaba's hosted flagship benched through Qoder, not the 27B model Clef is built on.

Do not read that 8.10 as a Clef score.

The closest local evidence I have for a 27B Qwen-family model is Qwable at 7.14.

The local stack I would build

Put Laya or Clef in the decision slot, answering typed questions about each job.

Send private, repetitive jobs to Qwable if you have the memory for a 27B model, or to a Gemma 4 build if you do not.

Keep one hosted lane, either Jev or a frontier model, for long option lists and hard reasoning.

That matters because on the 77-category test hosted Jev scored more than double what untrained Laya did.

Use Router(preload=True) with Laya if you see more than one language, because without it each language switch took around 20 seconds on my Mac.

Keep the decision endpoint as one config value, so swapping Laya, Clef and Jev never touches your agents.

Who should actually run Jev-style decisions locally

Run it locally if privacy is the whole point, such as client inboxes, payslips or anything regulated.

Run it locally if you make thousands of small decisions a day and want zero per-call cost.

Stay hosted if accuracy on long category lists matters more than privacy, because that is where the gap is widest.

And if you are not sure, start with the free hosted Jev door and move the private lanes local later.

Also On Our Network

🌐 a step-by-step local Jev setup

🌐 what running Jev locally means for your margins

🌐 the verdict on local Jev alternatives

🌐 how the local decision slot fits an Agent OS

🌐 Julian's hands-on Laya AI test

So the honest answer to "can I run Jev locally" is no, but a local decision model plus a benched local builder gets you the same shape of system today.

FAQ

Can I run Jev locally?

No. TypeSafe has not released Jev's weights, so Jev only runs as a hosted service. Its SDKs are open source, but they only call the hosted model.

What is the best local alternative to Jev?

Laya (pip install laya) is the one Julian installed and tested: about 20 ms per question, but untrained it scored 28 to 36 of 60 on his email test against Jev's 54. Cloudflare's Clef and Clef-flash have open weights and are described by Cloudflare as Jev-API compatible, but Julian has not tested them yet.

Which local model should do the work behind a local decision layer?

On GoldieBench the best local builder is Qwable 5 27B Coder at 7.14, followed by Agents-A1 at 4.83 and Gemma 4 12B Coder at 4.25.

Is Clef as good as Qwen 3.8 on GoldieBench?

There is no Clef score on GoldieBench. The Qwen 3.8 entry (8.10) is Alibaba's hosted flagship benched through Qoder, not the Qwen 3.8-27B model Clef is built on.

Does a local decision model still give confidence scores?

Laya does, and in Julian's test none of its 60 email answers reached 0.85 confidence, so the confidence number stayed honest even when the answers were wrong.

About Julian

I'm Julian Goldie — AI entrepreneur, SEO expert, and founder of the AI Profit Boardroom (4,000+ members).

I help business owners scale with AI agents, automation, and SEO.

400K+ YouTube subscribers. 7-figure AI agency (Goldie Agency). Daily training inside the Boardroom.

→ Get my best AI training inside the AI Profit Boardroom

Related reading

→ The Best Local Model For Hermes Agent — Decided By 45 Real Builds

→ The Best LLMs For Hermes Agent, Ranked By Real Work

→ The Best Free AI Model For Hermes Agent (All Three $0 Lanes)

🌐 Sister-site take: read this on agentos.guide

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly

📺 Video notes + links to the tools 👉 AI Profit Boardroom

🆓 Get a FREE AI Course + Community + 1,000 AI Agents 👉 AI Money Lab

🎥 Learn how I make these videos 👉 agentos.guide

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly