AI BENCHMARK PROFILE
LEVANTE-bench
LEVANTE-bench evaluates vision-language models on six cognitive tasks from the LEVANTE dataset, comparing model performance and error patterns with children aged 5-12 across three countries.
- Released
- 2026-06-03
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
LEVANTE-bench addresses the gap in evaluating VLMs against human cognitive development, providing a structured comparison that can inform model design and understanding of alignment with human cognition.
Motivation
Given the inherently multimodal nature of human experience, vision-language models (VLMs) hold substantial promise for modeling human cognition as it grows and develops with experience.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.