Benchmark Radar is a system created to address the challenges of finding relevant AI benchmarks and understanding their associated evaluations. It combines daily discovery of benchmark papers, repositories, datasets, and releases with a searchable benchmark catalog. The catalog currently contains 1,283 source records drawn from 4 benchmark catalogs and 12,916 numeric observations on 790 records. The system retains source identities and citations, allowing users to inspect benchmark evidence. Daily discovery is powered by 37 sources, including 13 direct connectors and 24 first-party research and engineering feeds.
The system includes a web dashboard with a benchmark leaderboard, Pareto frontier views, saturation and trend views, daily feeds, downloadable evidence, and a command-line interface (CLI) for offline queries. A worked example demonstrates how to use the catalog for a prior-art search, showing how to query the catalog and inspect benchmark evidence when designing a new evaluation. This allows engineers to assess the suitability of benchmarks for their specific model or agent applications.
Benchmark Radar’s goal is to reduce the time and effort required to locate and understand AI benchmarks. This is important for engineers who need to select appropriate benchmarks for evaluating model performance, tracking adoption trends, and identifying potential saturation issues within the benchmark landscape. The system’s comprehensive data and search capabilities contribute to more reliable and reproducible AI evaluations.
Source: https://arxiv.org/abs/2609.11115