AISE-Bench
AISE-Bench is a benchmark for evaluating multi-step API-using LLM agents in information seeking on academic knowledge graphs. It comprises 1,133 QA pairs with API trajectories, validated parameters, and grounded answers, and evaluates answer quality, reference grounding, API-planning correctness, and execution success.
- Released
- 2026-06-16
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Existing benchmarks for tool-using agents on academic graphs rely on synthetic or narrow tasks. AISE-Bench addresses this gap with real-world, full-cycle annotated data, enabling quantitative assessment of stepwise correctness, grounded summarization, and traceable reasoning in complex API workflows. It provides a challenging testbed for improving agent reliability in realistic information-seeking scenarios.
Motivation
Large language models (LLMs) augmented with tools are emerging as autonomous agents capable of using Web engine, APIs, and code to solve complex, long-horizon tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.