AI BENCHMARK PROFILE
CEO-Bench
CEO-Bench evaluates LLM agents on strategic resource allocation in multi-round organizational simulations with conflicting advisor inputs.
- Released
- 2026-06-16
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It probes executive decision-making capabilities beyond isolated cognitive tasks, revealing tradeoffs in agent behavior.
Motivation
Evaluating the decision-making capabilities of large language models (LLMs) is a growing research priority, yet existing benchmarks focus on isolated cognitive tasks such as reasoning, knowledge retrieval, and economic rationality in stylized settings.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.