GoldieBench deep dive
DeepSeek V4 Pro's full benchmark breakdown.
DeepSeek's flagship tier — benched head-to-head against its own cheap Flash. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.
01 · The headline numbers
No scored tasks yet - this model is on the bench but unranked.
02 · Every benchmark, bar by bar
All 0 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.
03 · Where it wins - with the judge's own words
04 · Where it struggles - quoted, not hidden
Boards that hide the weak rows are brochures. These are DeepSeek V4 Pro's lowest scored builds, verdicts unedited:
Nothing under 8.0 on the board right now - DeepSeek V4 Pro's floor is unusually high.
05 · Category breakdown vs the whole field
Games (0 tasks)
Others (0 tasks)
Pages (0 tasks)
Sims (0 tasks)
Visuals (0 tasks)
06 · Outside signals - every row sourced
My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.
No externally sourced rows tracked for this model yet - the GoldieBench scores above are the record.
07 · Head-to-head records
Not enough shared scored tasks yet.
08 · Methodology + honest limits
Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.
Frequently asked questions
01How much does DeepSeek V4 Pro cost?
DeepSeek's flagship tier, benched on `deepseek-v4-pro` via api.deepseek.com — the exact same 50 one-shot prompts, skill-infused game prompt and pipeline as the V4 Flash 0731 run, so the two runs are directly comparable side by side. DeepSeek's own line on the 0731 Flash refresh is that its post-training now beats the older V4-Pro-Preview — this run tests the current Pro against that claim.
Source ledger
- 01Official vendor site: api-docs.deepseek.comapi-docs.deepseek.com
Every model's breakdown
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.