Benchmark Radar
AI BENCHMARK PROFILE

TRAPSBench

General AIMultimodal Perception

TRAPSBench is a procedurally generated video benchmark with 1,404 physics pairs to test epistemic restraint in VLMs, using Penalized Epistemic Calibration Score (PECS).

Released
2026-08-13
Readiness
Paper only
Primary field
General AI

Why it matters

It highlights that VLMs can internally detect when abstention is required but fail to express it, offering a metric for calibration and restraint evaluation.

Motivation

When visual evidence is occluded or chaotic, models should abstain.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.