Benchmark Radar
AI BENCHMARK PROFILE

R3-Bench

General AICoding & Software Engineering

Evaluates resource-rational reasoning by measuring how models allocate a shared budget across six-problem suites in mathematics, competitive programming, and abstract reasoning, with scoring based on completed problems.

Released
2026-08-17
Readiness
Runnable
Primary field
General AI

Why it matters

Addresses the gap between demonstrated single-problem competence and shared-budget allocation, providing a fixed dataset and protocol for studying resource allocation in LLMs.

Motivation

In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.