3DCodeBench
3DCodeBench evaluates vision-language model agents on procedural 3D modeling by converting text and image references into Blender Python code. It includes 212 object categories with ground-truth scripts, and scores outputs on executability, image similarity (SigLIP-2/DINOv3), 3D shape distance (Chamfer/Uni3D), and LLM-as-judge, plus a human-preference ranking platform.
- Released
- 2026-05-31
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
There is no standardized way to compare model abilities in procedural 3D code generation. This benchmark provides a fixed protocol and dataset, enabling reproducible evaluation and comparison across models and coding-agent settings, useful for selecting models or guiding development of procedural modeling capabilities.
Motivation
Procedural 3D modeling through code is emerging as a versatile paradigm, offering deterministic, engine-ready, and precisely editable assets that neural 3D generators inherently lack.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.