Benchmark Radar
AI BENCHMARK PROFILE

3DCodeBench

General AIMultimodal PerceptionGoogle

3DCodeBench evaluates vision-language model agents on procedural 3D modeling by converting text and image references into Blender Python code. It includes 212 object categories with ground-truth scripts, and scores outputs on executability, image similarity (SigLIP-2/DINOv3), 3D shape distance (Chamfer/Uni3D), and LLM-as-judge, plus a human-preference ranking platform.

Released
2026-05-31
Readiness
Runnable
Primary field
General AI

Why it matters

There is no standardized way to compare model abilities in procedural 3D code generation. This benchmark provides a fixed protocol and dataset, enabling reproducible evaluation and comparison across models and coding-agent settings, useful for selecting models or guiding development of procedural modeling capabilities.

Motivation

Procedural 3D modeling through code is emerging as a versatile paradigm, offering deterministic, engine-ready, and precisely editable assets that neural 3D generators inherently lack.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.