1 paper
Alexander Rubinstein, Benjamin Raible, Martin Gubri +1
Evaluating modern machine learning models has become prohibitively expensive. Benchmarks such as LMMs-Eval and HELM demand thousands of GPU hours per model. Costly evaluation reduc…