AI BENCHMARK PROFILE
Harness-Bench
Harness-Bench evaluates configuration-level harness effects in agent workflows with 106 sandboxed tasks, measuring completion, process quality, efficiency, and failure behavior across model-harness pairings.
- Released
- 2026-05-27
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses a gap in agent evaluation by isolating harness configuration effects, showing that agent capability is configuration-level rather than model-only, and identifying execution-alignment failures.
Motivation
LLM agents are increasingly deployed as executable systems that use tools, modify workspaces, and produce concrete artifacts.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.