Benchmark Radar
AI BENCHMARK PROFILE

AuthMem-Bench

General AIKnowledge & Reasoning

AuthMem-Bench evaluates authority collapse in persistent memory for LLM agents. It uses a paired benchmark holding claims and tasks fixed while varying source authority, measuring write-time collapse, authorization errors, and automatic authority preservation.

Released
2026-08-03
Readiness
Paper only
Primary field
General AI

Why it matters

Memory consolidation can erase authority constraints, leading to unauthorized actions. AuthMem-Bench provides a controlled benchmark to measure and improve authority preservation in memory systems, relevant for safe agent deployment.

Motivation

Persistent memory allows (self-evolving) LLM agents to adapt across tasks by consolidating heterogeneous interaction histories into reusable facts, preferences, observations, and rules.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.