Benchmark · CodeCompetitive

CadEval

CadEval · evaluates the ability to generate and reason about Computer-Aided Design code, testing spatial reasoning and engineering knowledge.

Updated 2025-06-17
Models tested
17
Top score
74.0
o3
Median
34.0
min 12.0
Top-5 spread
σ 7.0
wide open

Best score over time · one chart, every benchmark

CADEVAL9 MODELS · FRONTIER RUNNING MAX0255075100SCORE ↑Nov 24Jan 25Mar 25Apr 25Jun 25RELEASE DATE →benchgecko.ai/benchmark/cadeval · frontier
Frontier on CadEval rose from 56.0 to 74.0 in 4 months · +18.0 points · latest leader o3 from OpenAI.
Pink dots = frontier records · 2 totalClick to open model page

Same category · related evaluations