Benchmark Radar
AI BENCHMARK PROFILE

MemSecBench

General AIKnowledge & Reasoning

Evaluates lifecycle security of agent memory systems with 310 cases across 48 contexts, using a Write-Execute-Forget protocol and evidence-based adjudication across seven checkpoints.

Released
2026-07-29
Readiness
Paper only
Primary field
General AI

Why it matters

Provides insight into how malicious instructions can persist in memory systems and affect later actions, highlighting security differences across configurations.

Motivation

Memory systems allow agents to retain and reuse information from past interactions, but they can also let malicious content persist.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.