Benchmark Radar
AI BENCHMARK PROFILE

RUT-Bench

General AIKnowledge & ReasoningMiaow-Lab

RUT-Bench evaluates LLM tool-use in realistic user interactions with 1,638 test samples across 59 executable environments, measuring success rate, informational honesty, and tool discipline.

Released
2026-06-02
Readiness
Runnable
Primary field
General AI

Why it matters

Fill the gap of real-world tool-calling evaluation by simulating non-ideal user behaviors, providing a standardized framework for assessing LLM robustness in practical scenarios.

Motivation

Despite great advances in tool-use capabilities of large language models (LLMs), existing evaluation benchmarks struggle to fully align with real-world scenarios.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.