AI BENCHMARK PROFILE
AllocBench
A paired benchmark tests whether LLM agents exhibit conscious allocation behavior under a fixed budget in an abstract text-based formulation and a code-construction task.
- Released
- 2026-07-25
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
It identifies a capability boundary in online tool allocation for frontier models, showing that abstract optimal behavior does not transfer to script-writing.
Motivation
Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.