AI BENCHMARK PROFILE
PhysElite
Evaluates multimodal LLMs on 11,586 bilingual Olympiad-level physics problems with visual diagrams, step-by-step solutions, and answer accuracy metrics.
- Released
- 2026-08-25
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Offers a large, expert-level multimodal physics benchmark with step-level process evaluation, addressing gaps in difficulty and visual coverage for physics reasoning.
Motivation
Understanding how (multimodal) large language models perform on physics problems requires benchmarks that reflect the difficulty and breadth of expert-level physical reasoning.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.