Benchmark Radar
AI BENCHMARK PROFILE

EIBench

General AIKnowledge & Reasoning

EIBench is a simulator-based benchmark for interactive emotion management, containing 2,222 scenarios across a 2x2 taxonomy of Support, Defense, Repair, and Charm. It evaluates LLM agents on multi-turn dialogue where a user simulator updates emotion-relation states and provides anchor-based scoring.

Released
2026-06-14
Readiness
Paper only
Primary field
General AI

Why it matters

EIBench fills the gap in evaluating emotional intelligence beyond static understanding, focusing on interactive emotion management over multiple turns. Its simulator provides both outcome and dense turn-level feedback, enabling training and evaluation in a unified environment.

Motivation

Emotional intelligence (EI) in Large Language Models (LLMs) is often evaluated through static understanding tasks or single-response dialogue generation.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.