Benchmark Radar
AI BENCHMARK PROFILE

CoffeeBench

Transport & LogisticsKnowledge & ReasoningSakana AI

CoffeeBench evaluates LLM agents as a coffee roaster in a 90-day multi-agent economy with fixed reference agents, measuring cumulative net income through autonomous communication, negotiation, and transactions.

Released
2026-06-15
Readiness
Runnable
Primary field
Transport & Logistics

Why it matters

Existing benchmarks often focus on single agents in static environments, whereas CoffeeBench captures long-horizon multi-agent economic interactions. It provides a scoring contract based on net income, enabling comparison of agent economic decision-making and revealing failure modes like idle-drift.

Motivation

As LLM agents become capable of increasingly long-horizon tasks, evaluating their performance in economic systems is becoming increasingly important.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.