1 paper
Daniel Simig, Tianlu Wang, Verna Dankers +4
In NLP, models are usually evaluated by reporting single-number performance scores on a number of readily available benchmarks, without much deeper analysis. Here, we argue that -…