Benchmark Radar
AI BENCHMARK PROFILE

PhysElite

General AIMultimodal PerceptionMathematics & Formal Sciences

Evaluates multimodal LLMs on 11,586 bilingual Olympiad-level physics problems with visual diagrams, step-by-step solutions, and answer accuracy metrics.

Released
2026-08-25
Readiness
Inspectable
Primary field
General AI

Why it matters

Offers a large, expert-level multimodal physics benchmark with step-level process evaluation, addressing gaps in difficulty and visual coverage for physics reasoning.

Motivation

Understanding how (multimodal) large language models perform on physics problems requires benchmarks that reflect the difficulty and breadth of expert-level physical reasoning.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.