HEART-Bench
HEART-Bench evaluates whether LLM agents can simulate coherent, human-like psychology. It provides 11 fictional characters with raw episodic memories, 64 decision-making scenarios based on the DIAMONDS taxonomy, and 673 multiple-choice questions with expert-annotated ground truth. Two evaluation tracks: MCQ and open-ended consciousness-narrative.
- Released
- 2026-05-28
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
This benchmark addresses the gap in evaluating emotional and personality consistency in LLM agents, complementing task-oriented benchmarks. It offers a standardized protocol for assessing psychological coherence, which is valuable for developing agents that can maintain stable personas and make value-consistent decisions in interactive applications.
Motivation
While LLM agents have demonstrated remarkable task-oriented abilities such as planning, reasoning, and action, few works have treated them as complete human personalities where emotional dimensions hold equal importance.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.