CAD Bench v2 is live

Benchmarks for AI Models and Agents on CAD Tasks

CAD Bench v2 evaluates models and AI agents on 100 held-out parametric FreeCAD tasks with a corrected, isolated verifier and public run receipts.

The v1 leaderboard is frozen and remains available as an archive. v2 is the current benchmark for all new runs.

Gear box CAD assembly
Parametric CAD Bench V2 Leaderboard
View full leaderboard →
RankModelAgentGeom ScoreSpec ScoreCombinedCost (USD)
1
Claude Fable 5.1(max)Claude Code84.6088.8484.81$198.58
2
GPT-6 Astra(max)Codex84.5189.1484.78$135.62
3
Grok 4.6(xhigh)Grok Build81.6987.7282.21$75.02
4
Claude Opus 5(max)Claude Code82.5086.1479.52$102.02
5
Kimi K3(max)mini-swe-agent77.4187.4876.18$42.98
6
Claude Sonnet 5(max)Claude Code71.5286.6470.52$227.55
7
GPT-5.6-Sol(max)Codex72.0479.8570.34$73.67
8
GPT-5.6 Terra(max)Codex71.0277.5568.66$64.53
9
Muse Spark 1.3(max)mini-swe-agent67.2678.8565.46$75.68
10
GLM-5.3(max)mini-swe-agent64.9285.7164.04$28.15
Benchmarks Tasks
Run Benchmark →

3D Parametric Part Generation

3D Parametric Part Generation

Generate editable, parametric 3D CAD part models from natural-language prompts and reference inputs.

Assembly Generation

Assembly Generation

Generate multi-part assemblies with proper mates, constraints, and component hierarchy.

Complex CAD workflow

Complex CAD workflow

Multi-step CAD workflows that generate, iteratively edit, and verify designs until the model meets the target spec.

Evaluation Methods
View Evaluator →

CAD evaluation should check not only visual similarity, but also whether the generated CAD is valid, accurate, rebuildable, and consistent with the design spec.

Each task is scored automatically in a sandboxed CAD environment, comparing the generated CAD against the design spec and reference CAD (ground truth) across the axes below. Scoring is deterministic — same output, same score.

Geometry Accuracy

Measures how closely the generated geometry matches the reference part.

Constraint & Assembly Correctness

Checks constraint satisfaction, mating validity, and assembly stability.

Parametric Correctness

Verifies CAD model consistency with the spec parameters.

Topology / Structure

Evaluates topological validity and part structure correctness.

Agent Workflow Success

Assesses task completion rate and workflow correctness.

Efficiency

Measures token usage, execution time, and resource efficiency.

Latest Updates

September 8, 2026

CAD Bench v2: A Corrected Benchmark for the New Frontier

CAD Bench v2 is now the current 100-task parametric FreeCAD benchmark. It upgrades the runtime to FreeCAD 1.1, isolates held-out grader assets, fixes reward reporting, and adds 10 fresh runs across Claude Fable 5.1, GPT-6 Astra, Grok 4.6, Kimi K3, Claude Opus 5, Claude Sonnet 5, GPT-5.6, Muse Spark 1.3, and GLM-5.3. The original v1 leaderboard is now frozen.


August 14, 2026

August 2026 Run: Opus 5 Leads at 0.906, Grok 4.6 good balance of cost and performance

Five new (agent, model) cells join the board, taking the top five spots; the leaderboard now holds 15 entries. Claude Opus 5 via Claude Code leads at 0.906 — the first result above 0.9 — closing 44% of the remaining gap to a perfect score while costing less than the previous leader. Grok 4.6 offers the best balance of cost and performance, delivering 97.6% of the top score at 67.6% of the cost. In the 91 days since the May run the engineering slice moved sharply — the weakest new cell would have ranked first in May, and the gain went almost entirely into geometry rather than spec.


May 13, 2026

Parametric CAD Bench Report

A new benchmark for AI agents that design parametric 3D mechanical parts in FreeCAD. Multi-step agentic loop scoring geometric correctness and consistency between the generated CAD and the provided spec, across 10 agent–model combinations spanning frontier vendors. Early results: GPT-5.5 via Codex leads at 0.832 with a visible harness effect — and is also the most expensive combination at $170.


May 10, 2026

Open-sourced freecad-validator on GitHub

A programmatic grader for evaluating AI-generated parametric FreeCAD parts — checks geometry similarity to the ground truth and CAD/spec consistency.


May 08, 2026

Published the cad-gen-freecad Dataset to Hugging Face

Open-sourced cad-gen-freecad on HuggingFace — native FreeCAD parts with parametric feature history, design specs, and renderings, aimed at accurate, parametric CAD generation from text prompt.

© 2025 gNucleus AI. All rights reserved.