Benchmark Radar
AI BENCHMARK PROFILE

WhisperBench

General AIKnowledge & Reasoning

WhisperBench evaluates stealth memory injection attacks on persistent personal agents through a 108-case benchmark spanning five risk categories with fact and preference poisoning, using an IMAP/SMTP workflow and an email agent skill.

Released
2026-07-06
Readiness
Paper only
Primary field
General AI

Why it matters

The benchmark addresses the evaluation gap in assessing security of persistent memory in AI agents, providing a way to measure susceptibility to memory injection attacks and the effectiveness of defenses.

Motivation

Persistent personal agents combine long-term memory with access to users' external environments, enabling personalized foreground assistance and proactive background execution.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.