Benchmark Radar
AI BENCHMARK PROFILE

Tailor-Bench

General AIMultimodal Perception

Tailor-Bench evaluates visual world models on simulating irregular physical interactions with three scenario modes (regular, unconventional, impossible) and predictive/descriptive generation settings.

Released
2026-06-23
Readiness
Runnable
Primary field
General AI

Why it matters

Current benchmarks focus on common interactions; Tailor-Bench exposes the long-tail gap in physical world modeling, testing generalization and constraint awareness beyond typical scenarios.

Motivation

Physical interactions follow a long-tailed distribution: a set of common and regular interactions dominates human experience and visual data, while a broad spectrum of rare and irregular interactions remains underrepresented.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.