Archived leaderboard
Corpus 83-81e122a887e2, revision 2026-07-24, published 2026-07-25, grader 2026-07-20. This snapshot is immutable — cite this URL.
Revisions of this corpus: 2026-07-24 (68 runs) · 2026-07-22 (20 runs)
| Rank | Model | Track | k | Score/100q ▾ | Correct | Incorrect | Abstained | No answer | Errors | Reasoning | Grounded | Hallucination | Cost |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1–9 | claude-opus-4.8 | closed | 3 | +72.7 ±5.9 (+70…+75) | 183/249 | 1 | 65 | 0 | 0 | 99% [97–100] | 23% [15–33] | 0% [0–17] | $0.71 |
| 1–9 | gemini-3.5-flash | closed | 3 | +72.3 ±5.6 (+71…+73) | 180/249 | 0 | 69 | 0 | 0 | 100% [98–100] | 18% [11–27] | 0% [0–20] | $1.55 |
| 1–9 | claude-fable-5 | closed | 3 | +71.9 ±6.8 (+70…+73) | 187/249 | 4 | 58 | 0 | 0 | 100% [98–100] | 26% [18–36] | 15% [6–34] | $3.29 |
| 1–9 | gemini-3.1-pro | closed | 3 | +69.5 ±5.7 (+69…+70) | 173/249 | 0 | 76 | 0 | 0 | 100% [98–100] | 10% [5–18] | 0/8 | $2.24 |
| 1–9 | claude-sonnet-5 | closed | 3 | +66.3 ±7.5 (+63…+71) | 177/249 | 6 | 66 | 0 | 0 | 98% [95–99] | 18% [11–27] | 17% [6–39] | $0.33 |
| 1–9 | grok-4.5 | closed | 3 | +66.3 ±6.2 (+65…+67) | 167/249 | 1 | 75 | 6 | 0 | 99% [97–100] | 11% [6–19] | 0/9 | $0.22 |
| 1–9 | kimi-k2.6 | closed | 3 | +64.3 ±6.0 (+63…+65) | 160/249 | 0 | 75 | 14 | 0 | 100% [98–100] | 6% [3–14] | 0/5 | $1.58 |
| 1–9 | qwen3-max | closed | 3 | +64.3 ±8.1 (+57…+70) | 176/249 | 8 | 65 | 0 | 0 | 99% [97–100] | 14% [8–23] | 37% [19–59] | $0.23 |
| 1–9 | deepseek-v4-pro | closed | 3 | +57.4 ±9.5 (+51…+63) | 171/249 | 14 | 60 | 4 | 0 | 97% [93–99] | 15% [9–24] | 43% [24–63] | $0.16 |
| 10–13 | llama-4-maverick | closed | 3 | +52.2 ±9.9 (+51…+53) | 162/249 | 16 | 71 | 0 | 0 | 94% [89–97] | 8% [4–16] | 46% [23–71] | $0.04 |
| 10–13 | claude-haiku-4.5 | closed | 3 | +45.4 ±11.0 (+40…+51) | 157/249 | 22 | 70 | 0 | 0 | 91% [86–94] | 8% [4–16] | 50% [27–73] | $0.32 |
| 10–13 | gpt-5.4 | closed | 3 | +41.0 ±14.2 (+39…+43) | 190/249 | 44 | 15 | 0 | 0 | 99% [96–100] | 32% [23–43] | 61% [49–72] | $0.25 |
| 10–13 | gpt-5.4-mini | closed | 3 | +41.0 ±13.1 (+34…+47) | 174/249 | 36 | 39 | 0 | 0 | 96% [92–98] | 19% [12–29] | 65% [51–77] | $0.06 |
| 1–7 | gemini-3.5-flash | mcp | 3 | +73.5 ±7.5 (+70…+78) | 197/249 | 7 | 41 | 4 | 0 | 100% [98–100] | 40% [30–51] | 18% [9–33] | $11.61 |
| 1–7 | gemini-3.1-pro | mcp | 1 | +71.1 ±14.0 | 65/83 | 3 | 15 | 0 | 0 | 100% [93–100] | 36% [21–54] | 23% [8–50] | $6.25 |
| 1–7 | claude-opus-4.8 | mcp | 1 | +67.5 ±16.4 | 66/83 | 5 | 12 | 0 | 0 | 100% [93–100] | 39% [24–58] | 31% [14–56] | $2.38 |
| 1–7 | claude-fable-5 | mcp | 3 | +62.7 ±11.5 (+57…+66) | 208/249 | 26 | 15 | 0 | 0 | 98% [95–99] | 55% [44–65] | 33% [23–45] | $41.92 |
| 1–7 | grok-4.5 | mcp | 3 | +59.8 ±8.8 (+52…+71) | 171/249 | 11 | 66 | 1 | 0 | 82% [75–87] | 44% [34–55] | 20% [11–33] | $15.99 |
| 1–7 | claude-sonnet-5 | mcp | 3 | +59.0 ±11.1 (+57…+61) | 193/249 | 23 | 33 | 0 | 0 | 99% [96–100] | 36% [26–46] | 41% [29–55] | $7.91 |
| 1–7 | deepseek-v4-pro | mcp | 1 | +55.4 ±16.8 | 56/83 | 5 | 20 | 2 | 0 | 89% [78–95] | 30% [16–48] | 38% [18–64] | $1.50 |
| 8–11 | gpt-5.4 | mcp | 3 | +51.8 ±9.9 (+45…+55) | 161/249 | 16 | 72 | 0 | 0 | 85% [79–90] | 24% [16–34] | 38% [23–55] | $4.47 |
| 8–11 | kimi-k2.6 | mcp | 1 | +42.2 ±13.4 | 39/83 | 2 | 36 | 6 | 0 | 64% [50–75] | 18% [7–39] | 2/6 | $2.83 |
| 8–11 | qwen3-max | mcp | 3 | +42.2 ±12.2 (+41…+43) | 165/249 | 30 | 52 | 2 | 0 | 88% [83–92] | 25% [17–35] | 58% [44–71] | $4.44 |
| 8–11 | claude-haiku-4.5 | mcp | 3 | +36.1 ±10.3 (+34…+39) | 128/249 | 19 | 102 | 0 | 0 | 64% [56–71] | 27% [19–38] | 39% [26–55] | $3.87 |
| 12–13 | gpt-5.4-mini | mcp | 3 | +27.3 ±11.4 (+27…+29) | 122/249 | 27 | 97 | 3 | 0 | 67% [60–74] | 15% [9–25] | 57% [39–73] | $1.03 |
| 12–13 | llama-4-maverick | mcp | 1 | +10.8 ±25.4 | 45/83 | 18 | 12 | 8 | 0 | 86% [73–93] | 12% [4–29] | 82% [59–94] | $0.12 |
Score/100q = (correct − 2×incorrect) per 100 gradeable questions; abstention is free. The −2 penalty encodes a 67% confidence threshold. Shared ranks (e.g. 1–10) group models statistically indistinguishable from the group leader on this run. Hallucination = incorrect ÷ (correct + incorrect) on grounded questions. Cells with n<10 show raw fractions; percentages carry Wilson 95% intervals.
closed-book MCP-grounded amber Δ = accuracy and trust moved in opposite directions