Benchmark Radar
AI BENCHMARK PROFILE

RetailBench

Transport & LogisticsKnowledge & Reasoning

RetailBench is a simulation benchmark evaluating tool-using LLM agents in single-store supermarket operations over a 180-day horizon. It covers pricing, replenishment, supplier selection, assortment, inventory aging, customer feedback, external events, and cash-flow constraints, with a privileged oracle policy for comparison.

Released
2026-06-14
Readiness
Paper only
Primary field
Transport & Logistics

Why it matters

RetailBench addresses the gap in evaluating LLM agents on long-horizon, economically grounded decision-making, where short-horizon tasks dominate existing benchmarks. It provides a controlled testbed for assessing coherent decision-making and reliability in dynamic environments.

Motivation

Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.