DungeonBench
DungeonBench evaluates tactical reasoning in Dungeons & Dragons combat with two tracks: Encounter for single fights and Day for linked encounters with persistent resources, using a shared decision stream of complete tactical observations and legal options.
- Released
- 2026-07-31
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Current benchmarks often under-test rules-rich tactical reasoning where geometry, timing, resources, and rule interactions matter. DungeonBench fills this gap by providing a reproducible simulator-based environment with clear scoring, allowing comparison of policies from heuristic controllers to language models.
Motivation
Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich tactical reasoning: the ability to choose well when geometry, timing, resources, objectives, and rule interactions all matter at once.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.