AI BENCHMARK PROFILE
HarnessRisk
HarnessRisk evaluates agent harness safety across six lifecycle phases with 128 sandboxed cases. It measures Utility, Attack Success Rate, Persistence, and Detection for each trajectory.
- Released
- 2026-08-18
- Readiness
- Runnable
- Primary field
- Cybersecurity
Why it matters
Addresses the need for systematic evaluation of agent harness safety across multiple responsibilities, enabling comparison of harness and model combinations and highlighting vulnerability patterns.
Motivation
Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and external actions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.