Benchmark Radar
AI BENCHMARK PROFILE

MMShopBench

General AIMultimodal Perception

MMShopBench is a real-log benchmark for multimodal multi-turn shopping agents. It uses cleaned shopping logs with annotations for purchase intent and mandatory product requirements. Agents must infer requirements from images and dialogue, retrieve candidates, and verify product satisfaction.

Released
2026-07-31
Readiness
Paper only
Primary field
General AI

Why it matters

Existing benchmarks often rely on text-only or synthetic requests, missing complex real-world multimodal shopping needs. MMShopBench provides a realistic evaluation to advance shopping agents, with an offline sandbox for reproducible experimentation.

Motivation

Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult to articulate in text alone.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.