WeClawArena
WeClawArena is a benchmark and runtime sandbox for multi-party owned-agent collaboration over personal workspaces. It contains 124 base tasks across six domains, expanded into 620 scenario variants with benign and attack-vector conditions, and reports task success and attack success separately.
- Released
- 2026-08-04
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing agent benchmarks do not provide an end-to-end sandbox for verifiable cross-user agent collaboration with realistic digital workspaces. WeClawArena enables evaluation of both collaborative task utility and security risks in human-centered agent networks, supporting diagnosis of privacy leakage, poisoned evidence, and invalid authority paths.
Motivation
Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and communicates with other agents through social and task relations.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.