Benchmark Radar
AI BENCHMARK PROFILE

UniClawBench

General AIKnowledge & ReasoningHKU-MMLab

A capability-driven benchmark for proactive agents in real-world tasks, with 400 bilingual tasks across five capabilities. It evaluates agents in Docker containers using step-by-step checkpoints and a closed-loop strategy with executor, supervisor, and user agents.

Released
2026-07-09
Readiness
Runnable
Primary field
General AI

Why it matters

Existing agent benchmarks rely on sandboxed environments and single-turn paradigms, which do not reflect real-world complexity. This benchmark provides a dynamic, capability-based evaluation that helps compare models and agent frameworks, aiding in identifying failure root causes.

Motivation

The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-world environments.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.