AI BENCHMARK PROFILE
Claw-Anything
Claw-Anything evaluates always-on LLM personal assistants in simulated environments with long-horizon activity histories, interdependent backend services, and GUI/CLI across devices. It includes 200 human-verified tasks scored on completion, robustness, communication, and safety, with a live leaderboard.
- Released
- 2026-05-25
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
It expands agent evaluation to broad, always-on contexts, revealing capability gaps in stateful, proactive assistance and supporting scalable data generation for training.
Motivation
Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.