DOE OSTI · 3019412
Ranking and Classifying AI Benchmarks
Abstract
We created a set of standards to efficiently evaluate AI benchmarks through objective means. Although prevalent, especially in recent times, AI benchmarks have no single way to measure their effectiveness. The MLCommons team provided a set of criteria for evaluating benchmarks, although the criteria lacks a clearly defined set of evaluation rules. We created a rubric with preset factors to efficiently and objectively evaluate a benchmark s quality. We created a software framework for processing lists of benchmarks for visualization. The framework and rating system allows researchers to quickly check if their benchmarks are effective.
Keep this discovery
Explore connections, maps & timelines
Shiraishi, Reece C. [Cornell U.], Hawks, Benjamin G. [Fermilab]. 2025-08-06. Ranking and Classifying AI Benchmarks. https://doi.org/10.2172/3019412
Cite the original work for its findings. Save a collection to share your selection of sources.