Benchmark Radar
AI BENCHMARK PROFILE

HUSH-Bench

General AIKnowledge & Reasoning

HUSH-Bench evaluates conversational agents' use of sensitive history under a conservative policy, with 2,400 prompts and matched no-memory references, measuring unsolicited integration and memory access.

Released
2026-06-04
Readiness
Paper only
Primary field
General AI

Why it matters

It isolates memory retention from use in dialogue generation, showing that retrieval systems may surface sensitive data even without user request. This motivates separating storage, retrieval, and scope decisions in memory design.

Motivation

Long-term memory helps conversational agents maintain continuity across sessions, while relevance and current-turn warrant remain distinct decisions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.