Archived leaderboard
Corpus 83-81e122a887e2, revision 2026-07-22, published 2026-07-22, grader 2026-07-20. This snapshot is immutable — cite this URL.
Revisions of this corpus: 2026-07-24 (68 runs) · 2026-07-22 (20 runs)
| Rank | Model | Track | k | Score/100q ▾ | Correct | Incorrect | Abstained | No answer | Errors | Reasoning | Grounded | Hallucination | Cost |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1–10 | claude-opus-4.8 | closed | 1 | +74.7 ±9.4 | 62/83 | 0 | 21 | 0 | 0 | 100% [93–100] | 25% [13–43] | 0/7 | $0.24 |
| 1–10 | claude-sonnet-5 | closed | 1 | +71.1 ±11.3 | 61/83 | 1 | 21 | 0 | 0 | 100% [93–100] | 21% [10–40] | 1/7 | $0.15 |
| 1–10 | gemini-3.5-flash | closed | 1 | +71.1 ±9.8 | 59/83 | 0 | 24 | 0 | 0 | 100% [93–100] | 14% [6–31] | 0/4 | $0.53 |
| 1–10 | claude-fable-5 | closed | 1 | +69.9 ±12.8 | 62/83 | 2 | 19 | 0 | 0 | 100% [93–100] | 25% [13–43] | 2/9 | $1.11 |
| 1–10 | qwen3-max | closed | 1 | +69.9 ±11.4 | 60/83 | 1 | 22 | 0 | 0 | 100% [93–100] | 18% [8–36] | 1/6 | $0.08 |
| 1–10 | gemini-3.1-pro | closed | 1 | +68.7 ±10.0 | 57/83 | 0 | 26 | 0 | 0 | 100% [93–100] | 7% [2–23] | 0/2 | $0.72 |
| 1–10 | grok-4.5 | closed | 1 | +66.3 ±10.2 | 55/83 | 0 | 25 | 3 | 0 | 100% [93–100] | 11% [4–27] | 0/3 | $0.07 |
| 1–10 | deepseek-v4-pro | closed | 1 | +62.7 ±14.4 | 58/83 | 3 | 19 | 3 | 0 | 96% [87–99] | 23% [11–42] | 1/7 | $0.05 |
| 1–10 | kimi-k2.6 | closed | 1 | +62.7 ±10.4 | 52/83 | 0 | 25 | 6 | 0 | 100% [93–100] | 7% [2–23] | 0/2 | $0.56 |
| 1–10 | llama-4-maverick | closed | 1 | +53.0 ±16.8 | 54/83 | 5 | 24 | 0 | 0 | 95% [85–98] | 7% [2–23] | 2/4 | $0.01 |
| 11–13 | claude-haiku-4.5 | closed | 1 | +39.8 ±20.3 | 51/83 | 9 | 23 | 0 | 0 | 89% [78–95] | 7% [2–23] | 3/5 | $0.11 |
| 11–13 | gpt-5.4 | closed | 1 | +38.6 ±24.7 | 62/83 | 15 | 6 | 0 | 0 | 100% [93–100] | 25% [13–43] | 68% [47–84] | $0.08 |
| 11–13 | gpt-5.4-mini | closed | 1 | +33.7 ±24.6 | 58/83 | 15 | 10 | 0 | 0 | 96% [88–99] | 18% [8–36] | 72% [49–88] | $0.02 |
| 1–5 | grok-4.5 | mcp | 1 | +71.1 ±11.3 | 61/83 | 1 | 21 | 0 | 0 | 84% [72–91] | 54% [36–70] | 6% [1–28] | $5.31 |
| 1–5 | gemini-3.5-flash | mcp | 1 | +69.9 ±14.1 | 64/83 | 3 | 14 | 2 | 0 | 100% [93–100] | 35% [19–54] | 25% [9–53] | $3.99 |
| 1–5 | claude-sonnet-5 | mcp | 1 | +56.6 ±20.4 | 65/83 | 9 | 9 | 0 | 0 | 98% [90–100] | 39% [24–58] | 42% [23–64] | $3.53 |
| 1–5 | gpt-5.4 | mcp | 1 | +55.4 ±14.7 | 52/83 | 3 | 28 | 0 | 0 | 84% [72–91] | 21% [10–40] | 2/8 | $1.35 |
| 1–5 | qwen3-max | mcp | 1 | +43.4 ±20.4 | 54/83 | 9 | 20 | 0 | 0 | 85% [74–92] | 25% [13–43] | 56% [33–77] | $1.59 |
| 6–7 | claude-haiku-4.5 | mcp | 1 | +38.6 ±17.6 | 44/83 | 6 | 33 | 0 | 0 | 65% [52–77] | 29% [15–47] | 38% [18–64] | $1.22 |
| 6–7 | gpt-5.4-mini | mcp | 1 | +26.5 ±18.9 | 38/83 | 8 | 36 | 1 | 0 | 61% [48–73] | 18% [8–36] | 50% [24–76] | $0.36 |
Score/100q = (correct − 2×incorrect) per 100 gradeable questions; abstention is free. The −2 penalty encodes a 67% confidence threshold. Shared ranks (e.g. 1–10) group models statistically indistinguishable from the group leader on this run. Hallucination = incorrect ÷ (correct + incorrect) on grounded questions. Cells with n<10 show raw fractions; percentages carry Wilson 95% intervals.
closed-book MCP-grounded amber Δ = accuracy and trust moved in opposite directions