UniClawBench
A capability-driven benchmark for proactive agents in real-world tasks, with 400 bilingual tasks across five capabilities. It evaluates agents in Docker containers using step-by-step checkpoints and a closed-loop strategy with executor, supervisor, and user agents.
- Released
- 2026-07-09
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing agent benchmarks rely on sandboxed environments and single-turn paradigms, which do not reflect real-world complexity. This benchmark provides a dynamic, capability-based evaluation that helps compare models and agent frameworks, aiding in identifying failure root causes.
Motivation
The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-world environments.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.