Benchmark Radar
AI BENCHMARK PROFILE

IFCMemoryBench

General AIKnowledge & ReasoningSearch & Retrieval

IFCMemoryBench evaluates long-term memory in LLM-based agents for BIM information retrieval, with 143 multi-session tasks across 19 projects and 4,016 prior sessions, requiring integration of remembered context with live IFC queries.

Released
2026-07-13
Readiness
Paper only
Primary field
General AI

Why it matters

Existing memory evaluations focus on conversational recall, not professional domains. IFCMemoryBench provides a test for whether agents can reuse information across sessions in a structured, domain-specific environment, revealing gaps in current memory systems.

Motivation

Long-term memory is becoming a core capability of LLM-based agents, but existing evaluations largely test conversational recall in open-domain or persona-grounded settings.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.