AI BENCHMARK PROFILE
D2VBench
D2VBench evaluates LLM value alignment using 10,000 daily dilemma scenarios covering 158 fine-grained value concepts, with a hybrid paradigm of multiple-choice and open-ended questions.
- Released
- 2026-07-22
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Value alignment benchmarks often lack scenario diversity; this provides a large, fine-grained dataset for assessing alignment across value dimensions in realistic settings.
Motivation
With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.