TriggerBench
TriggerBench is a benchmark for evaluating prospective memory (PM) in LLMs, spanning five dimensions across daily assistant and professional workflow scenarios, with matched retrospective memory (RM) controls, contrastive variants, and overloaded triggers, measuring proactive recall, false-alarm rate, and attentional robustness.
- Released
- 2026-06-22
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing LLM evaluations focus on retrospective memory via explicit queries, leaving prospective memory – the ability to spontaneously act on latent constraints – unevaluated. TriggerBench provides a granular measurement of PM capabilities, revealing a precision-recall trade-off, attentional fragility, and a decay with context length that RM does not exhibit, informing deployment decisions for long interactive applications.
Motivation
While Large Language Models (LLMs) are increasingly deployed in long interactions, existing evaluations focus predominantly on retrospective memory (RM) via explicit queries.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.