AI BENCHMARK PROFILE
GraphVerse
GraphVerse evaluates multimodal large language models on visual graph reasoning, covering perception, reasoning, and text-based graph reasoning in single and paired image settings, with process-sensitive scoring.
- Released
- 2026-08-07
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Provides a unified benchmark for visual graph reasoning that goes beyond answer-only metrics, addressing gaps in existing evaluations and enabling assessment of reasoning quality.
Motivation
Recent Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse vision-language tasks, creating an urgent need for more challenging benchmarks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.