UXBench
UXBench evaluates user experience in AI assistants through three tasks: UX Judge (binary classification of response quality), UX Eval (response generation), and UX Recovery (repairing failed interactions). The dataset contains 7,400 test instances from 70K+ real interaction logs, covering 8 scenarios and 83 domains. Scoring uses accuracy for Judge and GRM-rated quality for generation tasks.
- Released
- 2026-06-08
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
UXBench addresses the gap in evaluating AI assistants beyond raw capability, focusing on user-perceived utility and preference alignment. It provides a structured way to measure how well models understand and improve user experience, offering practical value for developing assistants that better satisfy real users.
Motivation
As AI assistants serve millions of users daily, evaluating user experience (UX) beyond general model capability has become increasingly important.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.