UserToolBench
UserToolBench evaluates personalized decision making in tool-use LLMs through inference of latent user preferences, clarification need, and user-aligned tool-call trajectories, built from privacy-sanitized interaction traces with 10 user profiles, 36 tool sets, and 1,065 turns.
- Released
- 2026-08-10
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Current personalization benchmarks focus on surface-level style imitation or response personalization, not whether models make correct decisions for the user. UserToolBench addresses this gap by emphasizing decision quality, which is critical for real-world delegation tasks.
Motivation
Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.