CoffeeBench
CoffeeBench evaluates LLM agents as a coffee roaster in a 90-day multi-agent economy with fixed reference agents, measuring cumulative net income through autonomous communication, negotiation, and transactions.
- Released
- 2026-06-15
- Readiness
- Runnable
- Primary field
- Transport & Logistics
Why it matters
Existing benchmarks often focus on single agents in static environments, whereas CoffeeBench captures long-horizon multi-agent economic interactions. It provides a scoring contract based on net income, enabling comparison of agent economic decision-making and revealing failure modes like idle-drift.
Motivation
As LLM agents become capable of increasingly long-horizon tasks, evaluating their performance in economic systems is becoming increasingly important.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.