SGR-Bench
SGR-Bench evaluates search agents on state-gated retrieval tasks. It includes 100 expert-curated tasks across 12 public data ecosystems, requiring agents to configure site-specific filters, views, hierarchies, or scopes to retrieve structured answers. Tasks come in goal-oriented and constraint-guided formulations.
- Released
- 2026-05-21
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
SGR-Bench addresses an undercharacterized class of retrieval tasks where evidence is hidden behind site-specific retrieval states. It provides a standardized evaluation to measure agents' ability to establish correct retrieval states, which is critical for real-world data retrieval from specialized websites.
Motivation
Recent advances in large language models and tool-using agents have expanded the range of benchmarked web tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.