AI BENCHMARK PROFILE
HERO'S JOURNEY
Tests rule induction in goal-directed episodic tasks through text games, covering eight tasks across attribute and procedural induction families with controllable lexical grounding and identifiability conditions.
- Released
- 2026-06-01
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Evaluates a specific cognitive capability in LLMs, revealing limitations in procedural induction and execution bottlenecks. Provides a structured testbed for studying rule learning and guiding method development.
Motivation
We introduce HERO'S JOURNEY, a benchmark for rule induction in goal-directed episodic tasks, where agents must infer hidden rules from demonstrations and act on them through multi-step execution.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.