Benchmark Radar
AI BENCHMARK PROFILE

OpenSciToolBench

General AIKnowledge & Reasoning

OpenSciToolBench is a benchmark with 900 tasks across four difficulty levels for evaluating LLM agents in open-world scientific tool acquisition.

Released
2026-07-30
Readiness
Paper only
Primary field
General AI

Why it matters

The benchmark supports a specific agent system's evaluation and lacks a standalone comparison path.

Motivation

Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computational tools.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.