Benchmark Radar
AI BENCHMARK PROFILE

HarnessSafe

General AISafety & Trustworthiness

HarnessSafe evaluates safety of agent harnesses through 328 executable cases across seven persistent-carrier families, using trace-based evaluation to track attack chains from entry to potential violation.

Released
2026-08-07
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the lack of benchmarks covering multiple persistent carriers and providing trace-level analysis, offering a more nuanced assessment of safety risks in agent systems compared to end-to-end success rates.

Motivation

Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills, tools, and shared artifacts.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.