Benchmark RadarRadar
GitHub

Build with us on GitHub

Explore both repositories.

benchmark-radar.com☆ View & star ↗ benchmark-radar.org☆ View & star ↗

A little star goes a long way. ♡

Benchmark Radar tracks emerging AI benchmarks.

Discover new public evaluations, search the benchmark library, and see which benchmarks are gaining attention or appearing in leading model reports.

Time window
Sort by
How Attention is calculated

What the score means. Attention is a relative 0–100 estimate of how much public interest a benchmark is receiving—or, for a brand-new release, may receive. It does not measure benchmark quality or model performance.

  • Latest: 45% Hugging Face paper votes + 25% GitHub stars + 5% dataset downloads, with up to 25 bonus points from the experimental LLM 7-day forecast.
  • 30 days: 30% paper votes + 55% GitHub stars + 15% dataset downloads. No LLM forecast.
  • 90 days: 15% paper votes + 55% GitHub stars + 30% dataset downloads. No LLM forecast.

How the forecast works. The LLM considers only the supplied benchmark scope, artifact readiness, topical breadth, and verified publisher evidence. It does not invent stars, votes, downloads, or adoption.

When data is missing. Each observed signal is converted to a percentile within the selected time window. A missing signal receives a neutral 50th-percentile value—not zero—and the result is marked with lower confidence.

Rising benchmarks