Benchmark Radar
AI BENCHMARK PROFILE

ChronoPhyBench

CybersecuritySafety & TrustworthinessChronoPhyBench Team

ChronoPhyBench evaluates multimodal LLMs on chronological physical dynamics reasoning via next-state prediction and VQA, using video frames and captions for single-image selection and multi-frame sorting.

Released
2026-06-06
Readiness
Paper only
Primary field
Cybersecurity

Why it matters

Addresses the gap of distinguishing true multimodal reasoning from language-prior exploitation, providing a robust framework to measure physical reasoning and hallucination rates in models.

Motivation

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in open-world reasoning and understanding.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.