AI BENCHMARK PROFILE
EduArt
EduArt evaluates art-historical knowledge and visual reasoning in multimodal LLMs using 871 human-authored questions in Italian and English, covering multiple formats and languages. Scoring is based on accuracy and psychometric properties.
- Released
- 2026-07-02
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
General benchmarks don't reveal discipline-specific capabilities. EduArt provides a fine-grained benchmark to assess art historical knowledge, showing that format significantly affects performance.
Motivation
Large language models now score near ceiling on general benchmarks, but these aggregate measures reveal little about how models behave within single disciplines.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.