Benchmark Radar
AI BENCHMARK PROFILE

Mat-Pref

General AIKnowledge & ReasoningMat-Pref team

Mat-Pref evaluates compositional reasoning in inorganic materials via 10,837 ionic-substitution questions across 11 structure families, with splits for in-distribution performance, held-out families, and cross-property transfer.

Released
2026-06-20
Readiness
Paper only
Primary field
General AI

Why it matters

It isolates generalization types (structural transfer, property transfer, memorization) in scientific reasoning, helping to identify when RLVR improves reasoning over memorization.

Motivation

Reinforcement learning from verifiable rewards (RLVR) has driven rapid progress in mathematical and code reasoning, but when extended to science, existing benchmarks do not decompose what generalizes: do gains reflect structural transfer, property transfer, or memorization?

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.