Benchmark Radar
AI BENCHMARK PROFILE

JarvisBench

General AIKnowledge & Reasoning

JarvisBench measures mediation in long-horizon agent workflows with two tracks: agent-collaboration and user-interaction. Built on WildClaw tasks and a reference Jarvis prototype.

Released
2026-07-18
Readiness
Inspectable
Primary field
General AI

Why it matters

Could fill a gap in evaluating agent-user interaction, but unclear path for reuse.

Motivation

Long-horizon AI agents are becoming increasingly capable, yet their interaction with users remains surprisingly thin.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.