Benchmark Radar
AI BENCHMARK PROFILE

EduArt

General AIMultimodal Perception

EduArt evaluates art-historical knowledge and visual reasoning in multimodal LLMs using 871 human-authored questions in Italian and English, covering multiple formats and languages. Scoring is based on accuracy and psychometric properties.

Released
2026-07-02
Readiness
Paper only
Primary field
General AI

Why it matters

General benchmarks don't reveal discipline-specific capabilities. EduArt provides a fine-grained benchmark to assess art historical knowledge, showing that format significantly affects performance.

Motivation

Large language models now score near ceiling on general benchmarks, but these aggregate measures reveal little about how models behave within single disciplines.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.