Benchmark Radar
AI BENCHMARK PROFILE

MuSR

General AIKnowledge & Reasoning

MuSR (Multistep Soft Reasoning) is a benchmark for evaluating language models on multistep soft reasoning tasks specified in natural language narratives. Created through a neurosymbolic synthetic-to-natural generation algorithm, it generates complex reasoning scenarios like murder mysteries roughly 1000 words in length that challenge current LLMs including GPT-4. The benchmark tests chain-of-thought reasoning capabilities across domains involving commonsense reasoning about physical and social situations.

Released
Unknown
Readiness
Paper only
Primary field
General AI

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.