Benchmark RadarRadar View on GitHub ↗

Official project information

About Benchmark Radar

Benchmark Radar is a daily-updated tracker and searchable library for public AI benchmarks, evaluation datasets, and the evaluations used in model reports.

What it tracks

Benchmark Radar connects new benchmark releases with their primary evidence: papers, code, datasets, public activity signals, and reported use by model developers.

How to use it

Use the Radar for recent releases, the Library for established evaluations, and Trends to examine how benchmark activity changes across AI capabilities and application fields.

Evidence policy

Records link back to public sources. Missing signals remain unknown rather than being converted to zero, and inclusion is not a quality endorsement.