AI BENCHMARK PROFILE
C-SUITEBENCH
C-SUITEBENCH evaluates multimodal LLMs as CEOs on five decision tasks under paired text-only and multimodal conditions across 50 scenarios, focusing on evidence-centric reasoning and constraint satisfaction.
- Released
- 2026-08-06
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing executive decision benchmarks are text-only, leaving unclear whether models can integrate visual evidence. C-SUITEBENCH reveals a multimodal integration paradox where visual inputs can degrade constrained allocation, informing selective grounding strategies.
Motivation
Large language models are increasingly applied as autonomous decision-making agents.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.