3 papers
cs.CL2024
Large-scale cloze evaluation reveals that token prediction tasks are neither lexically nor semantically aligned
Cassandra L. Jacobs, Loïc Grobol, Alvin Tsang
In this work we compare the generative behavior at the next token prediction level in several language models by comparing them to human productions in the cloze task. We find that…
cs.CL2024
Kreyòl-MT: Building MT for Latin American, Caribbean and Colonial African Creole Languages
Nathaniel R. Robinson, Raj Dabre, Ammon Shurtz +14
A majority of language technologies are tailored for a small number of high-resource languages, while relatively many low-resource languages are neglected. One such group, Creole l…
cs.CL2024
CreoleVal: Multilingual Multitask Benchmarks for Creoles
Heather Lent, Kushal Tatariya, Raj Dabre +18
Creoles represent an under-explored and marginalized group of languages, with few available resources for NLP research.While the genealogical ties between Creoles and a number of h…