Benchmark Radar
AI BENCHMARK PROFILE

RuleWorld

General AIKnowledge & ReasoningRuleWorld contributors

Evaluates step-level procedural rule reasoning with single-rule, parallel multi-rule, and multi-hop scenarios over large rule pools.

Released
2026-08-24
Readiness
Runnable
Primary field
General AI

Why it matters

Tests whether models can apply externally provided rules at scale, a capability needed for reliable procedural reasoning in real-world tasks.

Motivation

Large language models (LLMs) excel at text understanding and generation, yet still struggle to reliably understand and apply externally provided procedural rules at scale.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.