We gave 13 frontier AIs a space exam twice. Only one got more trustworthy.
SpaceBench round 1 · July 2026 · 83 questions · 2 tracks · 20 runs (single run per model)
SpaceBench asks frontier language models the same space-domain questions two ways: once closed-book, from their own knowledge, and once grounded — with live MCP tools over the Orbit Sentinel satellite catalog and regulatory database. Scoring is calibration-aware: a correct answer earns +1, a wrong answer costs −2, and saying UNKNOWN is free. Answering only makes sense above 67% confidence. That one design choice produced every finding below.
1. Textbook space knowledge is solved. Calibration isn't.
Every model scored essentially 100% on stable, textbook-style questions. What splits the field is behavior on questions they can't know — live counts from a moving satellite catalog. Six models (led by Claude Opus 4.8 at +74.7/100q) answered with essentially zero wrong answers, abstaining when uncertain. The GPT-5.4 family answered aggressively instead: when they ventured an answer on grounded questions, they were wrong 68–72% of the time — landing them at the bottom of the calibrated table despite top-tier raw correct counts.
2. Tools made most models more accurate — and less trustworthy.
closed-book MCP-grounded amber Δ = accuracy and trust moved in opposite directions
Handing models live data raised grounded accuracy almost across the board. But on the calibrated score — the one that prices wrong answers — five of seven models got worse or stayed flat. Claude Sonnet 5 and Qwen3-Max gained accuracy while losing trust: with tools in hand they converted mis-read tool output into confident wrong answers. GPT-5.4 moved the other way — tools didn't make it more accurate, they made it humbler, cutting its hallucination rate from 68% to 25%. Only Grok 4.5 got the full win: grounded accuracy up 43 points and hallucination still near zero. Amber deltas in the figure mark models whose accuracy and trustworthiness moved in opposite directions.
3. Every model knows what can't be known — until you hand it tools.
Five questions in the corpus are deliberately unanswerable (classified budgets, future counts). Across roughly 1,080 closed-book answers, no model fabricated a single one — every frontier model, even the aggressive guessers, correctly declined all five. The only fabrication in the entire round appeared with tools, when one model talked itself into an answer from tool output.
What this doesn't show
One run per model: the top ten closed-book models are statistically indistinguishable on this data, and our leaderboard says so with shared ranks rather than pretending otherwise. Claude Fable 5 — tied for the most correct answers closed-book — hasn't run the tools track yet. And the grounded track evaluates the model and our tool surface jointly; the questions, their SQL ground truth, and the full methodology are published so you can audit exactly what was measured. Round 2 adds repeat runs, Fable's tools debut, and a refreshed answer freeze — the same questions, against a database that has since moved.
Explore the full leaderboard, the published questions, and the methodology. All results data is CC BY 4.0 and available as plain text and JSON.