Benchmark Radar
AI BENCHMARK PROFILE

VisualLeakBench

General AIMultimodal Perception

VisualLeakBench is a 500-image benchmark for evaluating action-boundary propagation failures in vision-language agents, with stratified subsets and oracle diagnostics.

Released
2026-05-29
Readiness
Paper only
Primary field
General AI

Why it matters

Targets a specific safety failure mode in VLAs, but lacks clear public reuse path and scoring contract details.

Motivation

Vision-language agents increasingly consume screenshots, documents, and user interfaces before writing to memory, sending messages, or invoking external tools.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.