Benchmark Radar
AI BENCHMARK PROFILE

ResearchClawBench

General AIKnowledge & ReasoningInternScience

ResearchClawBench evaluates autonomous scientific research agents across 40 tasks from 10 domains, each grounded in a published paper with provided literature and raw data. It uses expert-curated multimodal rubrics to score target-paper-level re-discovery.

Released
2026-05-28
Readiness
Runnable
Primary field
General AI

Why it matters

Autonomous research agents claim to accelerate science, but their end-to-end capability is unverified. ResearchClawBench provides a standardized evaluation frontier to measure progress toward reliable research re-discovery.

Motivation

AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.