Benchmark Radar
AI BENCHMARK PROFILE

ClawArena-Team

General AIMultimodal Perception

ClawArena-Team evaluates a single text-only LLM's ability to manage a fixed, locally served pool of subagents (LLM, VLM, omni) across 41 multi-turn, multimodal, multi-directory scenarios with 258 evaluation rounds and 72 staged updates. Scoring is execution-based via shell commands, producing a Subagent-Management Score (SMS) that multiplies task correctness by a least-privilege and modality-routing factor, without LLM judges.

Released
2026-06-30
Readiness
Runnable
Primary field
General AI

Why it matters

Existing agent benchmarks measure a policy's own task-solving or emergent behavior of fixed multi-agent systems, but not the leadership capability of a single model orchestrating subagents. ClawArena-Team fills this gap by isolating management skill from raw capability, supporting decisions on model selection for delegation-heavy workflows and revealing cost-quality trade-offs.

Motivation

Production large language-model (LLM) agents are increasingly deployed not as lone problem-solvers but as managers: a main model creates specialized subagents, delegates work, and orchestrates their parallel, asynchronous returns through dynamic workflows.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.