Benchmark Radar
AI BENCHMARK PROFILE

ContextWeave

General AIKnowledge & Reasoning

ContextWeave is a longitudinal benchmark for evaluating memory in office workflows, with 1,005 executable tasks from 14 participants. It measures workspace quality and preference alignment.

Released
2026-08-05
Readiness
Paper only
Primary field
General AI

Why it matters

It addresses evaluation of memory in long-horizon agent workflows, but the lack of public artifacts and scoring details makes it non-reusable without further information.

Motivation

Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.