Benchmark Radar
AI BENCHMARK PROFILE

BIM-Edit

General AIKnowledge & Reasoning

BIM-Edit evaluates large language models on natural-language editing of Industry Foundation Classes (IFC) building models. The benchmark includes 324 editing tasks across 11 realistic building models and 36 synthetic scenes. Tasks are categorized as direct, spatial, or topological instructions, and outputs are scored on geometric accuracy, semantic validity, and topological consistency.

Released
2026-06-18
Readiness
Paper only
Primary field
General AI

Why it matters

Construction and architectural design rely on structured BIM models; the evaluation gap is that existing benchmarks mostly test geometry and creation from scratch. BIM-Edit measures scene understanding and semantic relational preservation, providing a capability signal for practical engineering workflows where editing is central.

Motivation

Large language models (LLMs) are increasingly applied to computer-aided design (CAD) to generate design artifacts from textual instructions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.