We ran the space benchmark three times. The winner changed.

SpaceBench round 2 · July 2026 · 83 questions · 3 runs per model · 68 runs pooled

Round 1 ran each model once and warned that its rankings were provisional — the top of the table was a statistical tie, and one lucky run could crown the wrong model. So we ran it again. Three times per model, same 83 questions, answers re-frozen from the live database first. Here's what changed, and what didn't.

1. The tools champion was luck.

What tools actually buy: trust vs. accuracy, closed-book → MCP-grounded
Trust — Score/100q Grounded accuracy — % 0 25 50 0 25 50 75 100 gpt-5.4 +11 -8pp gemini-3.1-pro +2 +26pp gemini-3.5-flash +1 +22pp deepseek-v4-pro -2 +15pp claude-opus-4.8 -5 +17pp grok-4.5 -6 +33pp claude-sonnet-5 -7 +18pp claude-haiku-4.5 -9 +19pp claude-fable-5 -9 +29pp gpt-5.4-mini -14 -4pp kimi-k2.6 -22 +12pp qwen3-max -22 +11pp llama-4-maverick -41 +3pp

closed-book MCP-grounded amber Δ = accuracy and trust moved in opposite directions

Round 1 named Grok 4.5 the tool-grounded winner on a single run of +43 points of grounded accuracy. Pooled over three runs, that run was its ceiling, not its center: Grok settles to third. The real leader is Gemini 3.5 Flash — the one model whose interval clears the pack, gaining accuracy with tools while barely denting its calibration. This is exactly the failure mode repeat runs exist to catch, and we'd rather show it happening than hide it.

2. Everything else replicated.

The calibration story from round 1 held precisely. Closed-book, the GPT-5.4 family again hallucinated on 60–65% of the grounded questions they attempted, while the abstainers stayed clean across all three runs. Run-to-run variance was small for most models — the numbers were stable; only the rankings between statistically tied models shuffled, which is the honest limit of a single run.

3. Tripling the data didn't break the closed-book tie.

With three times the evidence, the top of the closed-book table is still a nine-way statistical tie — Opus 4.8 leads, but eight models overlap it. That isn't a limitation to fix; it's a finding. On stable space knowledge, the frontier has converged, and what separates a good model from a bad one is no longer what it knows but whether it admits what it doesn't. With tools, that consensus breaks apart — and only one model moves decisively up.

What's still true, and still not shown

Claude Fable 5 made its tools debut here and lands second — the highest raw grounded accuracy on the board, taxed by the same Claude-family tendency to over-answer with tools in hand. Five models still have only a single MCP run (marked k=1, with wide intervals). The grounded track still measures the model and our tool surface together. The round-1 write-up is preserved unchanged as the record of what one run showed. Every number here is reproducible from the current leaderboard, the published questions, and the methodologyCC BY 4.0.