AI BENCHMARK PROFILE
JarvisBench
JarvisBench measures mediation in long-horizon agent workflows with two tracks: agent-collaboration and user-interaction. Built on WildClaw tasks and a reference Jarvis prototype.
- Released
- 2026-07-18
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Could fill a gap in evaluating agent-user interaction, but unclear path for reuse.
Motivation
Long-horizon AI agents are becoming increasingly capable, yet their interaction with users remains surprisingly thin.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.