AI BENCHMARK PROFILE
WorldBench
WorldBench evaluates multimodal large language models on visually diverse reasoning questions, with accuracy as the metric on a curated dataset.
- Released
- 2026-06-04
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Highlights visual diversity gaps in existing benchmarks; provides a challenging fixed dataset for model comparison.
Motivation
In real-world applications, models are expected to perform reliably across diverse settings.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.