Benchmark Radar
AI BENCHMARK PROFILE

DMV-Bench

General AIAgentsDMV-Bench Team

DMV-Bench evaluates visual memory of multimodal agents in an interactive shopping environment with 1,000 product variants. Agents run autonomous sessions and must recall cued product images via exact URL match, with text leakage controlled.

Released
2026-06-25
Readiness
Runnable
Primary field
General AI

Why it matters

Existing agent memory benchmarks focus on text; DMV-Bench isolates visual memory needs in interactive settings, offering a controlled protocol for comparing agent architectures on pixel-based recall across varying session lengths.

Motivation

Research on agent memory has matured rapidly, but almost entirely on the text side: few existing benchmarks ask, in an interactive environment, when an agent genuinely needs to remember what it saw rather than what it could write down.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.