Benchmark Radar
AI BENCHMARK PROFILE

D2VBench

General AIKnowledge & ReasoningTianjin University NLP Lab

D2VBench evaluates LLM value alignment using 10,000 daily dilemma scenarios covering 158 fine-grained value concepts, with a hybrid paradigm of multiple-choice and open-ended questions.

Released
2026-07-22
Readiness
Runnable
Primary field
General AI

Why it matters

Value alignment benchmarks often lack scenario diversity; this provides a large, fine-grained dataset for assessing alignment across value dimensions in realistic settings.

Motivation

With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.