1 paper
Arda Uzunoglu, Tianjian Li, Daniel Khashabi
Benchmarks are central to measuring progress in language models, but aggregate scores can obscure substantial variation across subdomains, making models appear broadly competent de…