H2HMem
H2HMem is a benchmark for evaluating memory capabilities of agents in human-human multimodal interactions, covering dyadic and multi-party conversations with tasks in memory recall, reasoning, and application.
- Released
- 2026-06-08
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
Existing memory benchmarks focus on single-user text interactions; H2HMem addresses the need for evaluating agents in complex multimodal human-human settings with asynchronous and conflicting information from multiple participants.
Motivation
Large language model agents are increasingly deployed in human-human interaction settings, such as meeting assistants and clinical documentation systems, where they must observe conversations and retain information for downstream queries.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.