What it tracks
Benchmark Radar connects new benchmark releases with their primary evidence: papers, code, datasets, public activity signals, and reported use by model developers.
Official project information
Benchmark Radar is a daily-updated tracker and searchable library for public AI benchmarks, evaluation datasets, and the evaluations used in model reports.
Benchmark Radar connects new benchmark releases with their primary evidence: papers, code, datasets, public activity signals, and reported use by model developers.
Use the Radar for recent releases, the Library for established evaluations, and Trends to examine how benchmark activity changes across AI capabilities and application fields.
Records link back to public sources. Missing signals remain unknown rather than being converted to zero, and inclusion is not a quality endorsement.
Benchmark Radar website
Official GitHub repository
Recent benchmark data
Benchmark library data