Benchmark Radar
AI BENCHMARK PROFILE

NCP-Bench

General AIAgents

NCP-Bench evaluates narrative commitment preservation in interactive narratives. It provides 100 movie-synopsis-based environments with structured narrative specifications and an automatic evaluator that checks fact, commitment, and trajectory consistency across narrator responses to player actions.

Released
2026-08-08
Readiness
Runnable
Primary field
General AI

Why it matters

The benchmark fills a gap in evaluating long-horizon logical consistency of LLM-driven interactive narrators, offering a repeatable protocol to compare models on narrative integrity under adversarial user interventions. This supports practical selection of models for interactive storytelling applications.

Motivation

The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.