Research · · 12 min read · Seldon Research

CADBench: Measuring computer-use agents in CAD

White concept sketch of a low sports car on a deep indigo field.

CADBench measures whether frontier agents can execute native, long-horizon mechanical-design work inside Autodesk Fusion and whether the resulting geometry, feature history and constraints survive deterministic verification.

tasks
105
unsolved tasks
68.1%
best pass rate
24.6%

Models

RankModelMean verifier score
01
OpenAIGPT-5.6 Sol
51.22
02
Google VertexGemini 3.7 Flash
50.15
03
xAIGrok 4.6
16.73
04
OpenAIGPT-5.6 Terra xhigh
16.38
05
MuseMuse Spark 1.2
13.44
06
QwenQwen 3.8 Max
8.76
07
Moonshot AIKimi K3
8.16
08
Inkling
6.75

Every CADBench task was authored by a human domain expert and grounded in a real CAD workflow from robotic-arm design to complex assembly. Every verifier is written by hand to model the underlying workflow. Scores combine intermediate process reward with outcome rewards, so a high score requires both sound construction and correct form-factor.

Harness choice

We ran Gemini 3.7 Flash with five harnesses to select the best one.

HarnessPrimary passMean verifier scoreReliabilityCostCost / taskCost / strict successTokens / taskTurns / task
ALE-Claw Selected24.6%50.1594.2%$55.76$0.53$2.162.45M155
OpenCode14.5%28.8792.8%$79.44$0.76$5.220.46M53
Antigravity7.3%23.0556.1%$290.05$2.76$37.841.94M220
Codex harness0.0%15.3278.3%$487.42$4.646.04M50
Prime Agent0.0%17.9878.1%$507.85$4.8435.83M224

Analysis

Higher score · lower cost0204060$0.00$2.18$4.35Cost per primary taskMean verifier scoreGPT-5.6 SolGemini 3.7Grok 4.6GPT-5.6 TerraMuse SparkQwen 3.8 MaxKimi K3InklingGPT-5.6 Sol — $3.93 · 51.22 mean verifier scoreGemini 3.7 Flash — $0.53 · 50.15 mean verifier scoreGrok 4.6 — $2.21 · 16.73 mean verifier scoreGPT-5.6 Terra xhigh — $3.41 · 16.38 mean verifier scoreMuse Spark 1.2 — ≈$0.67 · 13.44 mean verifier scoreQwen 3.8 Max — $3.19 · 8.76 mean verifier scoreKimi K3 — $2.82 · 8.16 mean verifier scoreInkling — $0.39 · 6.75 mean verifier score
ModelCostCost / taskCost / strict success
Gemini 3.7 Flash$55.76$0.53$2.16
GPT-5.6 Sol$413.12$3.93$16.97
Grok 4.6$232.28$2.21
Muse Spark$70.26$0.67
GPT-5.6 Terra$358.12$3.41
Kimi K3$295.65$2.82
Qwen 3.8 Max$774.77$7.38
Inkling$94.35$0.90

Task evidence

Switch between three tasks and inspect the recorded video and full provider-visible trace for each major model.

TASK
MODEL
MODELGPT-5.6 Sol · xhigh
VERIFIER SCORE23.24 / 100Strict pass: no
5.77Mreported tokens
29m 06swall time
TASK 1GPT-5.6 Sol · xhigh
VERIFIER SCORE23.24 / 100
SCREEN RECORDINGAutodesk Fusion · Windows 11

Methodology

COST LIMITS

500 turns per run

Every model is run three times per task. We cap the models at a maximum of 500 turns.

Stay posted if new models are added

We will email you when a new model is added to the CADBench results.