Benchmark Radar
AI BENCHMARK PROFILE

OmniPhys

General AIMultimodal PerceptionZJUKG

OmniPhys is a benchmark of 1,551 text-to-image generation samples grounded in a Physical Knowledge Graph, aligned with PhET simulations and curricula. It evaluates physical commonsense in generated images using a dual-path verification protocol.

Released
2026-07-28
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks use coarse descriptions and fail to diagnose specific physical principles. OmniPhys provides a fine-grained, curriculum-aligned evaluation to identify systemic physical reasoning gaps in image generation models, supporting targeted improvement.

Motivation

While text-to-image models exhibit remarkable visual fidelity, they frequently violate fundamental physical commonsense.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.