MPAR-Bench
MPAR-Bench evaluates multi-point associative reasoning in LLMs across English and Chinese using 1,000 items with diverse clues. Scoring includes exact-match accuracy, ANLS, embedding similarity, and reasoning-trace verification, with four perturbation types.
- Released
- 2026-08-11
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Current benchmarks focus on reasoning depth (longer chains) but neglect breadth (parallel semantic exploration). MPAR-Bench fills that gap, showing that depth does not guarantee robust breadth, and offers a practical tool for assessing models on a complementary reasoning dimension.
Motivation
Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.