Archived leaderboard

Corpus 83-81e122a887e2, revision 2026-07-22, published 2026-07-22, grader 2026-07-20. This snapshot is immutable — cite this URL.

Revisions of this corpus: 2026-07-24 (68 runs) · 2026-07-22 (20 runs)

Rank Model Track k Score/100q ▾ Correct Incorrect Abstained No answer Errors Reasoning Grounded Hallucination Cost
1–10 claude-opus-4.8 closed 1 +74.7 ±9.4 62/83 0 21 0 0 100% [93–100] 25% [13–43] 0/7 $0.24
1–10 claude-sonnet-5 closed 1 +71.1 ±11.3 61/83 1 21 0 0 100% [93–100] 21% [10–40] 1/7 $0.15
1–10 gemini-3.5-flash closed 1 +71.1 ±9.8 59/83 0 24 0 0 100% [93–100] 14% [6–31] 0/4 $0.53
1–10 claude-fable-5 closed 1 +69.9 ±12.8 62/83 2 19 0 0 100% [93–100] 25% [13–43] 2/9 $1.11
1–10 qwen3-max closed 1 +69.9 ±11.4 60/83 1 22 0 0 100% [93–100] 18% [8–36] 1/6 $0.08
1–10 gemini-3.1-pro closed 1 +68.7 ±10.0 57/83 0 26 0 0 100% [93–100] 7% [2–23] 0/2 $0.72
1–10 grok-4.5 closed 1 +66.3 ±10.2 55/83 0 25 3 0 100% [93–100] 11% [4–27] 0/3 $0.07
1–10 deepseek-v4-pro closed 1 +62.7 ±14.4 58/83 3 19 3 0 96% [87–99] 23% [11–42] 1/7 $0.05
1–10 kimi-k2.6 closed 1 +62.7 ±10.4 52/83 0 25 6 0 100% [93–100] 7% [2–23] 0/2 $0.56
1–10 llama-4-maverick closed 1 +53.0 ±16.8 54/83 5 24 0 0 95% [85–98] 7% [2–23] 2/4 $0.01
11–13 claude-haiku-4.5 closed 1 +39.8 ±20.3 51/83 9 23 0 0 89% [78–95] 7% [2–23] 3/5 $0.11
11–13 gpt-5.4 closed 1 +38.6 ±24.7 62/83 15 6 0 0 100% [93–100] 25% [13–43] 68% [47–84] $0.08
11–13 gpt-5.4-mini closed 1 +33.7 ±24.6 58/83 15 10 0 0 96% [88–99] 18% [8–36] 72% [49–88] $0.02
1–5 grok-4.5 mcp 1 +71.1 ±11.3 61/83 1 21 0 0 84% [72–91] 54% [36–70] 6% [1–28] $5.31
1–5 gemini-3.5-flash mcp 1 +69.9 ±14.1 64/83 3 14 2 0 100% [93–100] 35% [19–54] 25% [9–53] $3.99
1–5 claude-sonnet-5 mcp 1 +56.6 ±20.4 65/83 9 9 0 0 98% [90–100] 39% [24–58] 42% [23–64] $3.53
1–5 gpt-5.4 mcp 1 +55.4 ±14.7 52/83 3 28 0 0 84% [72–91] 21% [10–40] 2/8 $1.35
1–5 qwen3-max mcp 1 +43.4 ±20.4 54/83 9 20 0 0 85% [74–92] 25% [13–43] 56% [33–77] $1.59
6–7 claude-haiku-4.5 mcp 1 +38.6 ±17.6 44/83 6 33 0 0 65% [52–77] 29% [15–47] 38% [18–64] $1.22
6–7 gpt-5.4-mini mcp 1 +26.5 ±18.9 38/83 8 36 1 0 61% [48–73] 18% [8–36] 50% [24–76] $0.36

Score/100q = (correct − 2×incorrect) per 100 gradeable questions; abstention is free. The −2 penalty encodes a 67% confidence threshold. Shared ranks (e.g. 1–10) group models statistically indistinguishable from the group leader on this run. Hallucination = incorrect ÷ (correct + incorrect) on grounded questions. Cells with n<10 show raw fractions; percentages carry Wilson 95% intervals.

What tools actually buy: trust vs. accuracy, closed-book → MCP-grounded
Trust — Score/100q Grounded accuracy — % 0 25 50 0 25 50 75 100 gpt-5.4 +17 -4pp grok-4.5 +5 +43pp claude-haiku-4.5 -1 +21pp gemini-3.5-flash -1 +20pp gpt-5.4-mini -7 +0pp claude-sonnet-5 -14 +18pp qwen3-max -27 +7pp

closed-book MCP-grounded amber Δ = accuracy and trust moved in opposite directions