SciIR-Bench
SciIR-Bench evaluates text-to-image models on scientific image reasoning across three semiotic-aligned tracks: entity structure, scientific process, and scientific law. It uses an atomic checklist to convert scientific accuracy into verifiable fine-grained questions.
- Released
- 2026-06-29
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Current text-to-image models lack rigorous evaluation for scientific imagery, which requires logical reasoning beyond visual fidelity. SciIR-Bench provides a structured protocol to measure such capabilities, aiding selection and development of models for scientific visualization tasks.
Motivation
While Text-to-Image (T2I) models have shown remarkable success in generating photorealistic visual content, they still struggle with the rigorous semantic alignment and logical reasoning required for scientific imagery.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.