Benchmark Radar
AI BENCHMARK PROFILE

DailyReport

General AIKnowledge & Reasoning

DailyReport is a benchmark with 150 open-ended daily search tasks and 3,546 rubrics, evaluating search agents on information-seeking tasks. Tasks are decomposed into subtasks with cascade rubrics across dimensions, yielding interpretable scores and a user preference score.

Released
2026-06-11
Readiness
Runnable
Primary field
General AI

Why it matters

Existing search agent benchmarks often use specialized or artificial tasks with coarse scoring, limiting interpretability and real-world relevance. DailyReport provides a more user-centric evaluation protocol for daily search tasks, offering fine-grained, dimension-level scores to inform agent development.

Motivation

Search Agents (SAs) typically leverage large language models (LLMs) to support complex information-seeking tasks by autonomously exploring web sources and synthesizing information into comprehensive responses.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.