Benchmark Radar
AI BENCHMARK PROFILE

EComAgentBench

General AIKnowledge & Reasoning

EComAgentBench evaluates LLM-based shopping agents on 662 long-horizon product selection tasks built from real Amazon data. Each task requires uncovering hidden requirements spread across an explicit query, tool-gated profile, and scripted clarification, verifying candidates, and committing to one product within 100 tool calls. Scoring uses typed, source-tagged rubrics for each requirement.

Released
2026-06-16
Readiness
Paper only
Primary field
General AI

Why it matters

Existing shopping benchmarks reveal full intent upfront, failing to reflect real-world requirements that emerge over time. EComAgentBench measures agents' ability to handle long-horizon interactions, providing a reproducible foundation for comparing dependable shopping assistance.

Motivation

As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the query, recorded in a profile, or revealed only when the right question is asked.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.