5 citations · 6 across the 2 of their papers we have counts for
3 papers
Evaluate & Evaluation on the Hub: Better Best Practices for Data and Model Measurements
Leandro von Werra, Lewis Tunstall, Abhishek Thakur +16
Evaluation is a key part of machine learning (ML), yet there is a lack of support and tooling to enable its informed and systematic practice. We introduce Evaluate and Evaluation o…
The AI Index 2022 Annual Report
Daniel Zhang, Nestor Maslej, Erik Brynjolfsson +10
Welcome to the fifth edition of the AI Index Report! The latest edition includes data from a broad set of academic, private, and nonprofit organizations as well as more self-collec…
No News is Good News: A Critique of the One Billion Word Benchmark
Helen Ngo, João G. M. Araújo, Jeffrey Hui +1
The One Billion Word Benchmark is a dataset derived from the WMT 2011 News Crawl, commonly used to measure language modeling ability in natural language processing. We train models…