AI BENCHMARK PROFILE
ArbiGraph
ArbiGraph is a benchmark generator for evaluating tool-assisted language agents' context management via scalable task graphs with exact automatic verification.
- Released
- 2026-07-22
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Context management is critical for long reasoning workflows; this generator allows controlled variation of task complexity, but the public path is incomplete.
Motivation
We introduce ARBIGRAPH, a benchmark generator for evaluating whether tool-assisted language agents can retain, update, compose, and discard task-relevant context across extended reasoning workflows.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.