RetailBench
RetailBench is a simulation benchmark evaluating tool-using LLM agents in single-store supermarket operations over a 180-day horizon. It covers pricing, replenishment, supplier selection, assortment, inventory aging, customer feedback, external events, and cash-flow constraints, with a privileged oracle policy for comparison.
- Released
- 2026-06-14
- Readiness
- Paper only
- Primary field
- Transport & Logistics
Why it matters
RetailBench addresses the gap in evaluating LLM agents on long-horizon, economically grounded decision-making, where short-horizon tasks dominate existing benchmarks. It provides a controlled testbed for assessing coherent decision-making and reliability in dynamic environments.
Motivation
Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.