Benchmark Radar
AI BENCHMARK PROFILE

LEVANTE-bench

General AIKnowledge & ReasoningLEVANTE-bench Authors

LEVANTE-bench evaluates vision-language models on six cognitive tasks from the LEVANTE dataset, comparing model performance and error patterns with children aged 5-12 across three countries.

Released
2026-06-03
Readiness
Paper only
Primary field
General AI

Why it matters

LEVANTE-bench addresses the gap in evaluating VLMs against human cognitive development, providing a structured comparison that can inform model design and understanding of alignment with human cognition.

Motivation

Given the inherently multimodal nature of human experience, vision-language models (VLMs) hold substantial promise for modeling human cognition as it grows and develops with experience.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.