AI BENCHMARK PROFILE
EcoAgent-Bench
EcoAgent-Bench evaluates LLM agents' economic decision-making under budget constraints across 304 tasks in five families, with priced actions and metrics for micro accuracy and economic consistency.
- Released
- 2026-08-06
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing agent benchmarks measure task completion without considering cost-effectiveness. EcoAgent-Bench explicitly tests the trade-off between completion and resource use, providing a distinct evaluation dimension.
Motivation
Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.