AI BENCHMARK PROFILE
SoCRATES
SoCRATES is a benchmark for evaluating proactive LLM mediators in realistic, multi-domain testbeds. It contains 600 conflict scenarios across eight domains, probing five socio-cognitive adaptation axes, with topic-localized scoring.
- Released
- 2026-06-04
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Mediation evaluation requires realistic trajectories and topic-specific scoring. SoCRATES provides a structured testbed with socio-cognitive variations, enabling reliable comparison of LLM mediators.
Motivation
Evaluating LLM mediators remains challenging, as mediation unfolds as a real-time trajectory shaped by disputants' shifting emotions, intentions, and context.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.