Benchmark Radar
AI BENCHMARK PROFILE

SGR-Bench

General AIKnowledge & ReasoningSearch & RetrievalPKUAIWeb

SGR-Bench evaluates search agents on state-gated retrieval tasks. It includes 100 expert-curated tasks across 12 public data ecosystems, requiring agents to configure site-specific filters, views, hierarchies, or scopes to retrieve structured answers. Tasks come in goal-oriented and constraint-guided formulations.

Released
2026-05-21
Readiness
Inspectable
Primary field
General AI

Why it matters

SGR-Bench addresses an undercharacterized class of retrieval tasks where evidence is hidden behind site-specific retrieval states. It provides a standardized evaluation to measure agents' ability to establish correct retrieval states, which is critical for real-world data retrieval from specialized websites.

Motivation

Recent advances in large language models and tool-using agents have expanded the range of benchmarked web tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.