DMV-Bench
DMV-Bench evaluates visual memory of multimodal agents in an interactive shopping environment with 1,000 product variants. Agents run autonomous sessions and must recall cued product images via exact URL match, with text leakage controlled.
- Released
- 2026-06-25
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing agent memory benchmarks focus on text; DMV-Bench isolates visual memory needs in interactive settings, offering a controlled protocol for comparing agent architectures on pixel-based recall across varying session lengths.
Motivation
Research on agent memory has matured rapidly, but almost entirely on the text side: few existing benchmarks ask, in an interactive environment, when an agent genuinely needs to remember what it saw rather than what it could write down.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.