Benchmark Radar
AI BENCHMARK PROFILE

H2HMem

Health & Life SciencesMultimodal Perception

H2HMem is a benchmark for evaluating memory capabilities of agents in human-human multimodal interactions, covering dyadic and multi-party conversations with tasks in memory recall, reasoning, and application.

Released
2026-06-08
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Existing memory benchmarks focus on single-user text interactions; H2HMem addresses the need for evaluating agents in complex multimodal human-human settings with asynchronous and conflicting information from multiple participants.

Motivation

Large language model agents are increasingly deployed in human-human interaction settings, such as meeting assistants and clinical documentation systems, where they must observe conversations and retain information for downstream queries.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.