AI BENCHMARK PROFILE
Tailor-Bench
Tailor-Bench evaluates visual world models on simulating irregular physical interactions with three scenario modes (regular, unconventional, impossible) and predictive/descriptive generation settings.
- Released
- 2026-06-23
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Current benchmarks focus on common interactions; Tailor-Bench exposes the long-tail gap in physical world modeling, testing generalization and constraint awareness beyond typical scenarios.
Motivation
Physical interactions follow a long-tailed distribution: a set of common and regular interactions dominates human experience and visual data, while a broad spectrum of rare and irregular interactions remains underrepresented.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.