AI BENCHMARK PROFILE
Almieyar-Oryx-BloomBench
Evaluates vision-language models on six cognitive levels (Remember to Create) with bilingual English-Arabic image-question-answer tasks.
- Released
- 2026-06-04
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Provides a cognitively grounded benchmark to diagnose reasoning strengths/weaknesses across levels and languages, addressing gaps in existing piecemeal evaluations.
Motivation
Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and chart meaningful progress toward human-like multimodal intelligence.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.