Benchmark Radar
AI BENCHMARK PROFILE

WorldCupArena

General AIKnowledge & Reasoning

WorldCupArena evaluates language models and deep-research agents on football forecasting across multiple tasks including result, score, player events, statistics, and tournament outcomes. It uses a composite score and supports adding new schedules for future events.

Released
2026-07-20
Readiness
Runnable
Primary field
General AI

Why it matters

This benchmark provides a dynamic, real-world testbed for evaluating models on multi-source reasoning and structured prediction with ground truth on a fixed schedule. It offers practical value in comparing model performance against market and human baselines, revealing differences in detailed predictions beyond result accuracy.

Motivation

Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.