AI BENCHMARK PROFILE
RUT-Bench
RUT-Bench evaluates LLM tool-use in realistic user interactions with 1,638 test samples across 59 executable environments, measuring success rate, informational honesty, and tool discipline.
- Released
- 2026-06-02
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Fill the gap of real-world tool-calling evaluation by simulating non-ideal user behaviors, providing a standardized framework for assessing LLM robustness in practical scenarios.
Motivation
Despite great advances in tool-use capabilities of large language models (LLMs), existing evaluation benchmarks struggle to fully align with real-world scenarios.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.