2.7k citations · 2.7k across the 9 of their papers we have counts for
23 papers
The Validity of Evaluation Results: Assessing Concurrence Across Compositionality Benchmarks
Kaiser Sun, Adina Williams, Dieuwke Hupkes
NLP models have progressed drastically in recent years, according to numerous datasets proposed to evaluate performance. Questions remain, however, about how particular dataset des…
The Gender-GAP Pipeline: A Gender-Aware Polyglot Pipeline for Gender Characterisation in 55 Languages
Benjamin Muller, Belen Alastruey, Prangthip Hansanti +7
Gender biases in language generation systems are challenging to mitigate. One possible source for these biases is gender representation disparities in the training and evaluation d…
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin, Kevin Stone +65
In this work, we develop and release Llama 2, a collection of pretrained and fine-tuned large language models (LLMs) ranging in scale from 7 billion to 70 billion parameters. Our f…
Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus
Alex Warstadt, Leshem Choshen, Aaron Mueller +3
We present the call for papers for the BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus. This shared task is intended for participants with an i…
The Curious Case of Absolute Position Embeddings
Koustuv Sinha, Amirhossein Kazemnejad, Siva Reddy +3
Transformer language models encode the notion of word order using positional information. Most commonly, this positional information is represented by absolute position embeddings…
Hi, my name is Martha: Using names to measure and mitigate bias in generative dialogue models
Eric Michael Smith, Adina Williams
All AI models are susceptible to learning biases in data that they are trained on. For generative dialogue models, being trained on real human conversations containing unbalanced g…