10 citations · 10 across the 3 of their papers we have counts for
3 papers · 1 filter
Evaluating Alignment of Behavioral Dispositions in LLMs
Amir Taubenfeld, Zorik Gekhman, Lior Nezry +8
As LLMs integrate into our daily lives, understanding their behavior becomes essential. In this work, we focus on behavioral dispositionsthe underlying tendencies that shape res…
MiTTenS: A Dataset for Evaluating Gender Mistranslation
Kevin Robinson, Sneha Kudugunta, Romina Stella +2
Translation systems, including foundation models capable of translation, can produce errors that result in gender mistranslation, and such errors can be especially harmful. To meas…
MADLAD-400: A Multilingual And Document-Level Large Audited Dataset
Sneha Kudugunta, Isaac Caswell, Biao Zhang +8
We introduce MADLAD-400, a manually audited, general domain 3T token monolingual dataset based on CommonCrawl, spanning 419 languages. We discuss the limitations revealed by self-a…