Benchmark Radar
AI BENCHMARK PROFILE

SoCRATES

General AIKnowledge & ReasoningDISL-Lab

SoCRATES is a benchmark for evaluating proactive LLM mediators in realistic, multi-domain testbeds. It contains 600 conflict scenarios across eight domains, probing five socio-cognitive adaptation axes, with topic-localized scoring.

Released
2026-06-04
Readiness
Runnable
Primary field
General AI

Why it matters

Mediation evaluation requires realistic trajectories and topic-specific scoring. SoCRATES provides a structured testbed with socio-cognitive variations, enabling reliable comparison of LLM mediators.

Motivation

Evaluating LLM mediators remains challenging, as mediation unfolds as a real-time trajectory shaped by disputants' shifting emotions, intentions, and context.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.