Benchmark Radar
AI BENCHMARK PROFILE

Almieyar-Oryx-BloomBench

General AIMultimodal PerceptionQCRI

Evaluates vision-language models on six cognitive levels (Remember to Create) with bilingual English-Arabic image-question-answer tasks.

Released
2026-06-04
Readiness
Runnable
Primary field
General AI

Why it matters

Provides a cognitively grounded benchmark to diagnose reasoning strengths/weaknesses across levels and languages, addressing gaps in existing piecemeal evaluations.

Motivation

Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and chart meaningful progress toward human-like multimodal intelligence.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.