Benchmark Radar
AI BENCHMARK PROFILE

EcoAgent-Bench

General AIKnowledge & Reasoning

EcoAgent-Bench evaluates LLM agents' economic decision-making under budget constraints across 304 tasks in five families, with priced actions and metrics for micro accuracy and economic consistency.

Released
2026-08-06
Readiness
Paper only
Primary field
General AI

Why it matters

Existing agent benchmarks measure task completion without considering cost-effectiveness. EcoAgent-Bench explicitly tests the trade-off between completion and resource use, providing a distinct evaluation dimension.

Motivation

Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.