AI BENCHMARK PROFILE
HarnessSafe
HarnessSafe evaluates safety of agent harnesses through 328 executable cases across seven persistent-carrier families, using trace-based evaluation to track attack chains from entry to potential violation.
- Released
- 2026-08-07
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the lack of benchmarks covering multiple persistent carriers and providing trace-level analysis, offering a more nuanced assessment of safety risks in agent systems compared to end-to-end success rates.
Motivation
Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills, tools, and shared artifacts.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.