Benchmark Radar
AI BENCHMARK PROFILE

Claw-Anything

General AIKnowledge & ReasoningLiberCoders

Claw-Anything evaluates always-on LLM personal assistants in simulated environments with long-horizon activity histories, interdependent backend services, and GUI/CLI across devices. It includes 200 human-verified tasks scored on completion, robustness, communication, and safety, with a live leaderboard.

Released
2026-05-25
Readiness
Runnable
Primary field
General AI

Why it matters

It expands agent evaluation to broad, always-on contexts, revealing capability gaps in stateful, proactive assistance and supporting scalable data generation for training.

Motivation

Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.