AI BENCHMARK PROFILE
CollabBench
Evaluates collaborative ability of LLM agents in cooperative game environments with diverse player profiles and proactive engagement.
- Released
- 2026-06-04
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Could fill gap in grounded collaborative benchmarks, but lacks public evidence of evaluation protocol or artifacts.
Motivation
While LLM-based agents excel at individual tasks, effective collaboration with realistic human partners remains challenging.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.