Back to CAD-Bench

September 24, 2026

•

Leaderboard Update

Parametric CAD Bench V3: Claude Opus 5.5 Takes the Lead

Today we are adding two Claude Opus 5.5 results to the Parametric CAD Bench V3 leaderboard. Claude Code reaches 61.03%, while mini-swe-agent with the multimodal bridge reaches 57.96%.

The two new rows rank first and second on the current leaderboard. GPT-6 Astra follows at 56.87% and Claude Fable 5.1 at 56.75%. The confidence intervals of all four rows overlap, so this ordering should not be treated as a statistically decisive separation.

The updated V3 leaderboard

Overall score is the mean continuous reward across all 100 tasks. Failed and unscored trials count as zero. Scored excludes trials ending in an exception, while Perfect means a reward of exactly 1.0. The 95% confidence interval is calculated over all 100 per-task rewards. Every row links to its public Harbor run.

RankModelAgentEffortCreateCreate + EditImage-to-CADOverall (95% CI)ScoredPerfectCost (USD)Run
1
Claude Opus 5.5Claude Code
2.1.280
max50.8739.6984.6661.03% ± 7.5399/1009/100$573.94View run
2
Claude Opus 5.5mini-swe-agent
2.4.6
max42.7941.3881.7757.96% ± 7.8293/1008/100$498.28View run
3
GPT-6 AstraCodex
0.154.0
max52.2644.3769.7056.87% ± 6.85100/1005/100$317.81View run
4
Claude Fable 5.1Claude Code
2.1.270
max55.4836.7572.7156.75% ± 7.1898/1007/100$1,056.04View run
5
Claude Opus 5Claude Code
2.1.270
max49.4836.5262.6450.86% ± 7.4599/1006/100$1,012.21View run
6
Gemini 3.8 Flashmini-swe-agent
2.4.6
high47.9431.2231.1036.19% ± 6.9794/1004/100$271.85View run
7
Grok 4.7Grok Build
1.0.30
high45.5734.6326.9134.82% ± 6.64100/1003/100$670.22View run
8
GPT-5.6 SolCodex
0.154.0
max36.2324.7829.2029.98% ± 6.17100/1001/100$291.99View run
9
Grok 4.6Grok Build
1.0.30
high41.6133.4513.0027.72% ± 6.30100/1002/100$390.94View run
10
Gemini 3.8 FlashAntigravity
1.2.7
high43.2825.3217.4527.56% ± 6.72100/1004/100$196.74View run
11
Kimi K3mini-swe-agent
2.4.6
max39.1123.4915.3624.92% ± 6.4096/1001/100$314.55View run
12
Muse Spark 1.3mini-swe-agent
2.4.6
max33.6124.338.8420.92% ± 5.8199/1001/100$357.03View run
13
GPT-5.6 TerraCodex
0.154.0
max23.3223.238.1817.24% ± 5.8499/1000/100$215.12View run

Opus 5.5 improves substantially over Opus 5

Holding Claude Code constant, the overall score rises from 50.86% with Opus 5 to 61.03% with Opus 5.5, a paired improvement of 10.18 percentage points (paired 95% CI: 5.22 to 15.13). The largest change is image-to-CAD, which rises by 22.02 points, from 62.64% to 84.66%. Creation improves by 1.39 points and create-and-edit by 3.17 points.

The newer model is also less expensive in this run. Known API cost falls from $1,012.21 to $573.94, a 43.3% reduction, while token use falls from 946.0M to 552.4M, a 41.6% reduction.

The two agent results are close

With Claude Opus 5.5 held constant, Claude Code leads mini-swe-agent by 3.08 percentage points overall, but the paired 95% interval runs from -1.15 to 7.31 points. Claude Code wins 27 tasks, mini-swe-agent wins 28, and 45 tie. Claude Code is 8.07 points higher on creation and 2.90 points higher on image-to-CAD; mini-swe-agent is 1.68 points higher on create-and-edit.

Claude Code produces 99 scored trials versus 93 for mini-swe-agent. The mini-swe-agent run costs $75.66 less, at $498.28 rather than $573.94.

New image-to-CAD highs

Both Opus 5.5 rows achieve the highest image-to-CAD scores reported on the leaderboard so far. Claude Code reaches 84.66% and mini-swe-agent reaches 81.77%, compared with the previous high of 72.71% from Claude Fable 5.1. The result reinforces the importance of both model capability and the way an agent exposes engineering drawings to the model.

The audit counts every final exception as zero and includes superseded retry usage in cost and token totals. The Claude Code row has one final output-limit error; the mini-swe-agent row has seven final errors. Claude Code also invokes its Opus 5 fallback on two zero-scoring tasks after unsuccessful Opus 5.5 calls. Those fallback calls add no score benefit, and their $9.83 cost is included in the row total.

© 2025 gNucleus AI. All rights reserved.