Benchmark Radar
AI BENCHMARK PROFILE

GraphVerse

General AIMultimodal PerceptionGraphVerse Team

GraphVerse evaluates multimodal large language models on visual graph reasoning, covering perception, reasoning, and text-based graph reasoning in single and paired image settings, with process-sensitive scoring.

Released
2026-08-07
Readiness
Runnable
Primary field
General AI

Why it matters

Provides a unified benchmark for visual graph reasoning that goes beyond answer-only metrics, addressing gaps in existing evaluations and enabling assessment of reasoning quality.

Motivation

Recent Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse vision-language tasks, creating an urgent need for more challenging benchmarks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.