Benchmark Radar
AI BENCHMARK PROFILE

Graphwalks BFS 128K

General AIMathematics & Formal SciencesLong Context & Memory

A graph reasoning benchmark that evaluates language models' ability to perform breadth-first search (BFS) operations on graphs with context length over 128k tokens, testing long-context reasoning capabilities.

Released
Unknown
Readiness
Paper only
Primary field
General AI

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.