Benchmark Radar
AI BENCHMARK PROFILE

HarnessRisk

CybersecuritySafety & Trustworthiness

HarnessRisk evaluates agent harness safety across six lifecycle phases with 128 sandboxed cases. It measures Utility, Attack Success Rate, Persistence, and Detection for each trajectory.

Released
2026-08-18
Readiness
Runnable
Primary field
Cybersecurity

Why it matters

Addresses the need for systematic evaluation of agent harness safety across multiple responsibilities, enabling comparison of harness and model combinations and highlighting vulnerability patterns.

Motivation

Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and external actions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.