EComAgentBench
EComAgentBench evaluates LLM-based shopping agents on 662 long-horizon product selection tasks built from real Amazon data. Each task requires uncovering hidden requirements spread across an explicit query, tool-gated profile, and scripted clarification, verifying candidates, and committing to one product within 100 tool calls. Scoring uses typed, source-tagged rubrics for each requirement.
- Released
- 2026-06-16
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing shopping benchmarks reveal full intent upfront, failing to reflect real-world requirements that emerge over time. EComAgentBench measures agents' ability to handle long-horizon interactions, providing a reproducible foundation for comparing dependable shopping assistance.
Motivation
As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the query, recorded in a profile, or revealed only when the right question is asked.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.