Back to CAD-Bench
Parametric CAD Bench V2 Leaderboard
CurrentHeld-out results on the corrected v2 task and verifier stack. Score is the mean continuous task reward; failed or unscored trials count as zero.
| Rank | Model | Agent | Effort | Score (95% CI) | Scored | Perfect | Avg sec | Tokens | Known cost | Date | Run |
|---|---|---|---|---|---|---|---|---|---|---|---|
1 | Claude Fable 5.1 | Claude Code 2.1.251 | max | 84.81% ± 4.13 | 100/100 | 46/100 | 415 | 34.06M | $198.58 | 2026-09-05 | View run |
2 | GPT-6 Astra | Codex 0.153.4 | max | 84.78% ± 4.17 | 100/100 | 45/100 | 227 | 43.02M | $135.62 | 2026-09-05 | View run |
3 | Grok 4.6 | Grok Build 1.0.5 | xhigh | 82.21% ± 4.77 | 100/100 | 47/100 | 452 | 77.27M | $75.02 | 2026-09-05 | View run |
4 | Claude Opus 5 | Claude Code 2.1.248 | max | 79.52% ± 5.69 | 100/100 | 49/100 | 376 | 47.72M | $102.02 | 2026-09-04 | View run |
5 | Kimi K3 | mini-swe-agent 2.4.3 | max | 76.18% ± 6.39 | 100/100 | 46/100 | 538 | 32.73M | $42.98 | 2026-09-07 | View run |
6 | Claude Sonnet 5 | Claude Code 2.1.251 | max | 70.52% ± 7.58 | 100/100 | 49/100 | 1222 | 461.42M | $227.55 | 2026-09-06 | View run |
7 | GPT-5.6-Sol | Codex 0.150.0 | max | 70.34% ± 7.01 | 100/100 | 43/100 | 226 | 73.57M | $73.67 | 2026-09-04 | View run |
8 | GPT-5.6 Terra | Codex 0.153.4 | max | 68.66% ± 7.19 | 100/100 | 47/100 | 327 | 137.55M | $64.53 | 2026-09-06 | View run |
9 | Muse Spark 1.3 | mini-swe-agent 2.4.3 | max | 65.46% ± 8.03 | 100/100 | 47/100 | 972 | 176.25M | $75.68 | 2026-09-07 | View run |
10 | GLM-5.3 | mini-swe-agent 2.4.3 | max | 64.04% ± 8.61 | 99/100 | 51/100 | 878 | 50.36M | $28.15 | 2026-09-07 | View run |