Benchmark Radar
AI BENCHMARK PROFILE

MafiaScope

General AIKnowledge & ReasoningMafiaScope Team

MafiaScope evaluates LLM agents in the social deduction game Mafia. It probes each agent's private beliefs after every public utterance, scoring them against ground truth without influencing the game. The testbed provides an open-source engine, interactive visualizer, recorded games, and counterfactual replay corpus.

Released
2026-07-12
Readiness
Runnable
Primary field
General AI

Why it matters

Machine Theory of Mind is difficult to measure from observable behavior alone. MafiaScope separates incorrect belief formation from incorrect action under correct beliefs, a distinction invisible in dialogue or outcome data, enabling more precise evaluation of social reasoning in LLM agents.

Motivation

An LLM agent's public behaviour reveals little about its social reasoning: an agent that votes correctly may be guessing, and an agent that lies well leaves no trace of what it actually believes.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.