AI BENCHMARK PROFILE
RuleWeaver
Evaluates rule-centered scenario reasoning through corpus-derived IF-THEN rules, composed into QA instances with rubric-based answer quality, rule recall, and rule precision scoring.
- Released
- 2026-08-27
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Provides process-level evaluation beyond final-answer correctness, exposing where models fail to apply complex rules in domain scenarios.
Motivation
Large language models (LLMs) are increasingly applied to specialized domains, where effective use of domain expertise often requires reasoning over complex rules in concrete scenarios.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.