NCP-Bench
NCP-Bench evaluates narrative commitment preservation in interactive narratives. It provides 100 movie-synopsis-based environments with structured narrative specifications and an automatic evaluator that checks fact, commitment, and trajectory consistency across narrator responses to player actions.
- Released
- 2026-08-08
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
The benchmark fills a gap in evaluating long-horizon logical consistency of LLM-driven interactive narrators, offering a repeatable protocol to compare models on narrative integrity under adversarial user interventions. This supports practical selection of models for interactive storytelling applications.
Motivation
The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.