AI BENCHMARK PROFILE
OmniPhys
Evaluates multimodal physics understanding, reasoning, and generation on 15,246 questions with 19,850 images from middle-school to university levels.
- Released
- 2026-08-26
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Fills the gap in comprehensive physics benchmarks for MLLMs, including structured diagram generation and fine-grained reasoning analysis.
Motivation
Multimodal Large Language Models (MLLMs) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.