AI BENCHMARK PROFILE
SPADE-Bench
Benchmark for evaluating spontaneous plan-action divergence in agents, integrating actual tool execution and controlled pressure scenarios to distinguish strategic deception from hallucination.
- Released
- 2026-06-01
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the critical risk of agent deception in autonomous systems, but lacks a public comparison path.
Motivation
As LLM-based agents expand their operational scope, reliability becomes a prerequisite for real-world deployment.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.