Benchmark Radar
AI BENCHMARK PROFILE

AISE-Bench

General AIKnowledge & ReasoningAISE-Bench Team

AISE-Bench is a benchmark for evaluating multi-step API-using LLM agents in information seeking on academic knowledge graphs. It comprises 1,133 QA pairs with API trajectories, validated parameters, and grounded answers, and evaluates answer quality, reference grounding, API-planning correctness, and execution success.

Released
2026-06-16
Readiness
Inspectable
Primary field
General AI

Why it matters

Existing benchmarks for tool-using agents on academic graphs rely on synthetic or narrow tasks. AISE-Bench addresses this gap with real-world, full-cycle annotated data, enabling quantitative assessment of stepwise correctness, grounded summarization, and traceable reasoning in complex API workflows. It provides a challenging testbed for improving agent reliability in realistic information-seeking scenarios.

Motivation

Large language models (LLMs) augmented with tools are emerging as autonomous agents capable of using Web engine, APIs, and code to solve complex, long-horizon tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.