AI BENCHMARK PROFILE
Momento
Momento benchmarks persistent agentic task completion in multi-session service environments, requiring agents to resolve temporal dependencies and evolving user goals across sessions.
- Released
- 2026-05-30
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Highlights the gap in agent evaluation by focusing on multi-session history and misestimation of user state, which is critical for realistic human-agent interaction.
Motivation
Recent advances in agentic AI have enabled agents to complete complex tasks through tool use, reasoning, and multi-step planning.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.