Benchmark Radar
AI BENCHMARK PROFILE

GABench

General AIKnowledge & Reasoning

GABench evaluates LLM agents on graph analysis tasks across three graph types and four task categories: graph retrieval, graph theory, graph machine learning, and graph open-ended QA. It provides 84 executable tools and 10,400 tasks with verifiable ground truth.

Released
2026-08-03
Readiness
Paper only
Primary field
General AI

Why it matters

Existing graph benchmarks lack coverage and typically format tasks as text QA, limiting agent evaluation. GABench offers a comprehensive, tool-based benchmark for assessing end-to-end agentic capabilities in graph analysis, providing practical insights into harness and tool-call quality.

Motivation

Large language model (LLM) agents are increasingly capable of planning, using tools, and interacting with external environments.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.