AI BENCHMARK PROFILE
RuleWorld
Evaluates step-level procedural rule reasoning with single-rule, parallel multi-rule, and multi-hop scenarios over large rule pools.
- Released
- 2026-08-24
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Tests whether models can apply externally provided rules at scale, a capability needed for reliable procedural reasoning in real-world tasks.
Motivation
Large language models (LLMs) excel at text understanding and generation, yet still struggle to reliably understand and apply externally provided procedural rules at scale.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.