EIBench
EIBench is a simulator-based benchmark for interactive emotion management, containing 2,222 scenarios across a 2x2 taxonomy of Support, Defense, Repair, and Charm. It evaluates LLM agents on multi-turn dialogue where a user simulator updates emotion-relation states and provides anchor-based scoring.
- Released
- 2026-06-14
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
EIBench fills the gap in evaluating emotional intelligence beyond static understanding, focusing on interactive emotion management over multiple turns. Its simulator provides both outcome and dense turn-level feedback, enabling training and evaluation in a unified environment.
Motivation
Emotional intelligence (EI) in Large Language Models (LLMs) is often evaluated through static understanding tasks or single-response dialogue generation.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.