Benchmark Radar
AI BENCHMARK PROFILE

ArbiGraph

General AIKnowledge & Reasoning

ArbiGraph is a benchmark generator for evaluating tool-assisted language agents' context management via scalable task graphs with exact automatic verification.

Released
2026-07-22
Readiness
Runnable
Primary field
General AI

Why it matters

Context management is critical for long reasoning workflows; this generator allows controlled variation of task complexity, but the public path is incomplete.

Motivation

We introduce ARBIGRAPH, a benchmark generator for evaluating whether tool-assisted language agents can retain, update, compose, and discard task-relevant context across extended reasoning workflows.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.