activity
20212024
most citedBetween words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

106 citations · 181 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL20243 cited

Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning

Shivalika Singh, Freddie Vargus, Daniel Dsouza +30

Datasets are foundational to many breakthroughs in modern artificial intelligence. Many recent achievements in the space of natural language processing (NLP) can be attributed to t…

cs.CL2024

CIDAR: Culturally Relevant Instruction Dataset For Arabic

Zaid Alyafeai, Khalid Almubarak, Ahmed Ashraf +9

Instruction tuning has emerged as a prominent methodology for teaching Large Language Models (LLMs) to follow instructions. However, current instruction datasets predominantly cate…

cs.CL20234 cited

Ashaar: Automatic Analysis and Generation of Arabic Poetry Using Deep Learning Approaches

Zaid Alyafeai, Maged S. Al-Shaibani, Moataz Ahmed

Poetry holds immense significance within the cultural and traditional fabric of any nation. It serves as a vehicle for poets to articulate their emotions, preserve customs, and con…

cs.CL202365 cited

The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

Hugo Laurençon, Lucile Saulnier, Thomas Wang +51

As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…

cs.CL20223 cited

Masader Plus: A New Interface for Exploring +500 Arabic NLP Datasets

Yousef Altaher, Ali Fadel, Mazen Alotaibi +18

Masader (Alyafeai et al., 2021) created a metadata structure to be used for cataloguing Arabic NLP datasets. However, developing an easy way to explore such a catalogue is a challe…

cs.CL2021106 cited

Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Sabrina J. Mielke, Zaid Alyafeai, Elizabeth Salesky +8

What are the units of text that we want to model? From bytes to multi-word expressions, text can be analyzed and generated at many granularities. Until recently, most natural langu…