Benchmark Radar
AI BENCHMARK PROFILE

CBX-Bench

General AIMultimodal PerceptionCBX-Bench Team

CBX-Bench is a benchmark for quantitatively measuring the quality of Concept Bottleneck Model (CBM) explanations. It uses a council of five open-weight multimodal LLMs to score explanations given an image and class, validated against human preferences. The benchmark maintains a leaderboard for CBM explanation quality.

Released
2026-08-15
Readiness
Runnable
Primary field
General AI

Why it matters

CBM interpretability is often evaluated by downstream accuracy, lacking quantitative measures of explanation quality. CBX-Bench offers a human-aligned, scalable evaluation that does not require concept ground truth, enabling comparison of explanation quality across different CBMs.

Motivation

Concept Bottleneck Models (CBMs) are designed to make visual classification interpretable by expressing predictions through human-understandable concepts.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.