Benchmark Radar
AI BENCHMARK PROFILE

CEO-Bench

General AIAgents

CEO-Bench evaluates LLM agents on strategic resource allocation in multi-round organizational simulations with conflicting advisor inputs.

Released
2026-06-16
Readiness
Paper only
Primary field
General AI

Why it matters

It probes executive decision-making capabilities beyond isolated cognitive tasks, revealing tradeoffs in agent behavior.

Motivation

Evaluating the decision-making capabilities of large language models (LLMs) is a growing research priority, yet existing benchmarks focus on isolated cognitive tasks such as reasoning, knowledge retrieval, and economic rationality in stylized settings.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.