AI BENCHMARK PROFILE
ContextWeave
ContextWeave is a longitudinal benchmark for evaluating memory in office workflows, with 1,005 executable tasks from 14 participants. It measures workspace quality and preference alignment.
- Released
- 2026-08-05
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It addresses evaluation of memory in long-horizon agent workflows, but the lack of public artifacts and scoring details makes it non-reusable without further information.
Motivation
Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.