activity
20152023
most citedExploring Nearest Neighbor Approaches for Image Captioning

162 citations · 594 across the 11 of their papers we have counts for

collaborators
Showing 2022Show all

5 papers · 1 filter

cs.AI2022★ 8 cited

Measuring Data

Margaret Mitchell, Alexandra Sasha Luccioni, Nathan Lambert +7

We identify the task of measuring data to quantitatively characterize the composition of machine learning data and datasets. Similar to an object's height, width, and volume, data…

cs.CL2022★ 41 cited

The Stack: 3 TB of permissively licensed source code

Denis Kocetkov, Raymond Li, Loubna Ben Allal +10

Large Language Models (LLMs) play an ever-increasing role in the field of Artificial Intelligence (AI)--not only for natural language processing but also for code understanding and…

cs.CL2022

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

BigScience Workshop, :, Teven Le Scao +391

Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to wi…

cs.CL2022★ 1 cited

SEAL : Interactive Tool for Systematic Error Analysis and Labeling

Nazneen Rajani, Weixin Liang, Lingjiao Chen +2

With the advent of Transformers, large language models (LLMs) have saturated well-known NLP benchmarks and leaderboards with high aggregate performance. However, many times these m…

cs.LG2022★ 5 cited

Evaluate & Evaluation on the Hub: Better Best Practices for Data and Model Measurements

Leandro von Werra, Lewis Tunstall, Abhishek Thakur +16

Evaluation is a key part of machine learning (ML), yet there is a lack of support and tooling to enable its informed and systematic practice. We introduce Evaluate and Evaluation o…