Benchmark Radar
AI BENCHMARK PROFILE

Relay-Bench

General AIKnowledge & Reasoning

Relay-Bench evaluates LLMs on multi-domain reasoning chains, presenting composite problems that combine subproblems from domains like visual reasoning, coding, math, information extraction, and data analysis. It measures overall accuracy on these complex, text-only tasks.

Released
2026-07-20
Readiness
Paper only
Primary field
General AI

Why it matters

The benchmark addresses the need for holistic evaluation of LLMs on combined reasoning across multiple domains, which is common in real-world tasks. It provides practical value in assessing models' ability to handle complex, multi-step problems, with current models scoring below 50%.

Motivation

Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability to complete an assortment of tasks from distinct domains in a single prompt.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.