Benchmark Radar
AI BENCHMARK PROFILE

MPAR-Bench

General AIKnowledge & ReasoningNo official publisher identified

MPAR-Bench evaluates multi-point associative reasoning in LLMs across English and Chinese using 1,000 items with diverse clues. Scoring includes exact-match accuracy, ANLS, embedding similarity, and reasoning-trace verification, with four perturbation types.

Released
2026-08-11
Readiness
Paper only
Primary field
General AI

Why it matters

Current benchmarks focus on reasoning depth (longer chains) but neglect breadth (parallel semantic exploration). MPAR-Bench fills that gap, showing that depth does not guarantee robust breadth, and offers a practical tool for assessing models on a complementary reasoning dimension.

Motivation

Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.