MMShopBench
MMShopBench is a real-log benchmark for multimodal multi-turn shopping agents. It uses cleaned shopping logs with annotations for purchase intent and mandatory product requirements. Agents must infer requirements from images and dialogue, retrieve candidates, and verify product satisfaction.
- Released
- 2026-07-31
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing benchmarks often rely on text-only or synthetic requests, missing complex real-world multimodal shopping needs. MMShopBench provides a realistic evaluation to advance shopping agents, with an offline sandbox for reproducible experimentation.
Motivation
Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult to articulate in text alone.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.