AI BENCHMARK PROFILE
HUSH-Bench
HUSH-Bench evaluates conversational agents' use of sensitive history under a conservative policy, with 2,400 prompts and matched no-memory references, measuring unsolicited integration and memory access.
- Released
- 2026-06-04
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It isolates memory retention from use in dialogue generation, showing that retrieval systems may surface sensitive data even without user request. This motivates separating storage, retrieval, and scope decisions in memory design.
Motivation
Long-term memory helps conversational agents maintain continuity across sessions, while relevance and current-turn warrant remain distinct decisions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.