Benchmark Radar
AI BENCHMARK PROFILE

UXBench

General AIKnowledge & Reasoning

UXBench evaluates user experience in AI assistants through three tasks: UX Judge (binary classification of response quality), UX Eval (response generation), and UX Recovery (repairing failed interactions). The dataset contains 7,400 test instances from 70K+ real interaction logs, covering 8 scenarios and 83 domains. Scoring uses accuracy for Judge and GRM-rated quality for generation tasks.

Released
2026-06-08
Readiness
Runnable
Primary field
General AI

Why it matters

UXBench addresses the gap in evaluating AI assistants beyond raw capability, focusing on user-perceived utility and preference alignment. It provides a structured way to measure how well models understand and improve user experience, offering practical value for developing assistants that better satisfy real users.

Motivation

As AI assistants serve millions of users daily, evaluating user experience (UX) beyond general model capability has become increasingly important.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.