Benchmark Radar
AI BENCHMARK PROFILE

SciExplore

General AIRobotics & Embodied IntelligenceSciExplore Team

SciExplore evaluates scientific information-seeking and reasoning in LLMs and agents with four task types: scientific database navigation, ambiguous literature retrieval, missing reference completion, and cross-source structured knowledge synthesis, spanning 103 expert-curated tasks across more than ten scientific disciplines.

Released
2026-07-23
Readiness
Paper only
Primary field
General AI

Why it matters

Existing benchmarks emphasize general-domain retrieval or static QA, leaving a gap in assessing realistic scientific workflows. SciExplore's progressive task complexity provides a way to compare model capabilities in entity-level reasoning, document identification, evidence grounding, and domain-level synthesis, informing deployment in scientific research contexts.

Motivation

Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.