Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)

GoldieBench deep dive

DeepSeek V4 Flash's full benchmark breakdown.

DeepSeek's cheap tier, retrained for agents — same size, sharper loops. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.

Data refreshed
2026-07-31
Scored tasks
0
Reading time
4 min
External sources
1

01 · The headline numbers

No scored tasks yet - this model is on the bench but unranked.

02 · Every benchmark, bar by bar

All 0 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.

03 · Where it wins - with the judge's own words

04 · Where it struggles - quoted, not hidden

Boards that hide the weak rows are brochures. These are DeepSeek V4 Flash's lowest scored builds, verdicts unedited:

Nothing under 8.0 on the board right now - DeepSeek V4 Flash's floor is unusually high.

05 · Category breakdown vs the whole field

Games (0 tasks)

field avg7.2

Others (0 tasks)

field avg7.6

Pages (0 tasks)

field avg7.6

Sims (0 tasks)

field avg6.8

Visuals (0 tasks)

field avg6.9

06 · Outside signals - every row sourced

My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.

No externally sourced rows tracked for this model yet - the GoldieBench scores above are the record.

07 · Head-to-head records

Not enough shared scored tasks yet.

08 · Methodology + honest limits

Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.

Frequently asked questions

01How much does DeepSeek V4 Flash cost?

Benched on the DeepSeek-V4-Flash-0731 public beta, launched 2026-07-31 on DeepSeek's official API. DeepSeek describe it as a major upgrade to agent capabilities whose benchmark scores now surpass their previous V4-Pro-Preview, using the exact same model architecture and size as the preview — the gain is post-training, not scale. Natively supports the Responses API format and is adapted for Codex-style coding loops. Benched against api.deepseek.co

Source ledger

  1. 01Official vendor site: api-docs.deepseek.comapi-docs.deepseek.com

Every model's breakdown

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly