arXiv stat.ML
· Papers
Benchmarking on Tasks That Matter: Dataset Selection for Preserving Model Rankings
arXiv:2606.27997v1 Announce Type: cross Abstract: Benchmarks of machine learning models often include many datasets, making evaluation expensive. For efficiency, it is preferable to perform evaluations on small, representative datasets instead. The selection of such subsets typically relies on heuristics and is rarely