Benchmark Radar
AI BENCHMARK PROFILE

SciFigBench

General AIKnowledge & Reasoning

SciFigBench is a diagnostic VLM benchmark for scientific figure understanding, covering perception, reasoning, and behavioral reliability under uncertainty.

Released
2026-08-13
Readiness
Inspectable
Primary field
General AI

Why it matters

It addresses the gap in evaluating VLMs' behavior when visual evidence is missing or misleading, critical for scientific workflows.

Motivation

Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading).

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.