Benchmark Radar
AI BENCHMARK PROFILE

MUSE

General AIKnowledge & ReasoningMUSE Benchmark Team

MUSE evaluates text-to-CAD generation of complex B-Rep assemblies via design specifications. It scores models on code validity, geometric correctness, and design-intent alignment using rubrics covering functionality, manufacturability, and assemblability. A VLM judge with human validation is used for scalable scoring.

Released
2026-05-27
Readiness
Inspectable
Primary field
General AI

Why it matters

Existing CAD benchmarks focus on single-part geometric similarity, missing industrial requirements. MUSE provides a structured evaluation that measures practical design quality, enabling progress toward engineering-ready CAD generation.

Motivation

Large language models (LLMs) have recently advanced text-driven 3D generation, yet Text-to-CAD remains far from supporting industrial product design.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.