AI BENCHMARK PROFILE
NarrativeWorldBench
NarrativeWorldBench evaluates long-horizon narrative generation in audio drama across nine structural metrics and four Indic languages, across horizons from 10 to 200 episodes. Scoring is based on plot-beat F1 and other metrics.
- Released
- 2026-06-16
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Current LLMs degrade on long-horizon narrative coherence. This benchmark provides metrics for evaluating long-form structured content generation and cross-lingual capabilities.
Motivation
Long-form serialized audio drama, with arcs that run for 200 to 800 episodes, is a major creative medium and a setting where frontier large language models (LLMs) fail.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.