ClawArena-Team
ClawArena-Team evaluates a single text-only LLM's ability to manage a fixed, locally served pool of subagents (LLM, VLM, omni) across 41 multi-turn, multimodal, multi-directory scenarios with 258 evaluation rounds and 72 staged updates. Scoring is execution-based via shell commands, producing a Subagent-Management Score (SMS) that multiplies task correctness by a least-privilege and modality-routing factor, without LLM judges.
- Released
- 2026-06-30
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing agent benchmarks measure a policy's own task-solving or emergent behavior of fixed multi-agent systems, but not the leadership capability of a single model orchestrating subagents. ClawArena-Team fills this gap by isolating management skill from raw capability, supporting decisions on model selection for delegation-heavy workflows and revealing cost-quality trade-offs.
Motivation
Production large language-model (LLM) agents are increasingly deployed not as lone problem-solvers but as managers: a main model creates specialized subagents, delegates work, and orchestrates their parallel, asynchronous returns through dynamic workflows.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.