Relay-Bench
Relay-Bench evaluates LLMs on multi-domain reasoning chains, presenting composite problems that combine subproblems from domains like visual reasoning, coding, math, information extraction, and data analysis. It measures overall accuracy on these complex, text-only tasks.
- Released
- 2026-07-20
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
The benchmark addresses the need for holistic evaluation of LLMs on combined reasoning across multiple domains, which is common in real-world tasks. It provides practical value in assessing models' ability to handle complex, multi-step problems, with current models scoring below 50%.
Motivation
Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability to complete an assortment of tasks from distinct domains in a single prompt.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.