Benchmark Radar
AI BENCHMARK PROFILE

UserToolBench

General AIKnowledge & Reasoning

UserToolBench evaluates personalized decision making in tool-use LLMs through inference of latent user preferences, clarification need, and user-aligned tool-call trajectories, built from privacy-sanitized interaction traces with 10 user profiles, 36 tool sets, and 1,065 turns.

Released
2026-08-10
Readiness
Paper only
Primary field
General AI

Why it matters

Current personalization benchmarks focus on surface-level style imitation or response personalization, not whether models make correct decisions for the user. UserToolBench addresses this gap by emphasizing decision quality, which is critical for real-world delegation tasks.

Motivation

Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.