Benchmark Radar
AI BENCHMARK PROFILE

WSE-bench

General AIKnowledge & Reasoning

WSE-bench evaluates LLM storytelling in open-ended world simulations, assessing three capacities: Generation Coverage (proportion of planned narrative steps), Consistency (canon coherence), and Richness (meaningful development). Uses a process benchmark to compare agent architectures.

Released
2026-08-16
Readiness
Paper only
Primary field
General AI

Why it matters

Storytelling evaluation has focused on finished stories, but open-ended narratives require sustained generation, coherence, and development. WSE-bench makes these dynamics visible and shows they are distinct capacities.

Motivation

Large language models can write fluent stories, but open-ended storytelling requires more than local fluency.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.