AI BENCHMARK PROFILE
SciFigBench
SciFigBench is a diagnostic VLM benchmark for scientific figure understanding, covering perception, reasoning, and behavioral reliability under uncertainty.
- Released
- 2026-08-13
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
It addresses the gap in evaluating VLMs' behavior when visual evidence is missing or misleading, critical for scientific workflows.
Motivation
Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading).
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.