activity
20202026
most citedMultitask Prompted Training Enables Zero-Shot Task Generalization

563 citations · 869 across the 25 of their papers we have counts for

collaborators
Showing 2021Show all

5 papers · 1 filter

cs.CL2021★ 106 cited

Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Sabrina J. Mielke, Zaid Alyafeai, Elizabeth Salesky +8

What are the units of text that we want to model? From bytes to multi-word expressions, text can be analyzed and generated at many granularities. Until recently, most natural langu…

cs.CL2021

Masader: Metadata Sourcing for Arabic Text and Speech Data Resources

Zaid Alyafeai, Maraim Masoud, Mustafa Ghaleb +1

The NLP pipeline has evolved dramatically in the last few years. The first step in the pipeline is to find suitable annotated datasets to evaluate the tasks we are trying to solve.…

cs.LG2021★ 563 cited

Multitask Prompted Training Enables Zero-Shot Task Generalization

Victor Sanh, Albert Webson, Colin Raffel +38

Large language models have recently been shown to attain reasonable zero-shot generalization on a diverse set of tasks (Brown et al., 2020). It has been hypothesized that this is a…

cs.CL2021

Calliar: An Online Handwritten Dataset for Arabic Calligraphy

Zaid Alyafeai, Maged S. Al-shaibani, Mustafa Ghaleb +1

Calligraphy is an essential part of the Arabic heritage and culture. It has been used in the past for the decoration of houses and mosques. Usually, such calligraphy is designed ma…

cs.CL2021★ 3 cited

Evaluating Various Tokenizers for Arabic Text Classification

Zaid Alyafeai, Maged S. Al-shaibani, Mustafa Ghaleb +1

The first step in any NLP pipeline is to split the text into individual tokens. The most obvious and straightforward approach is to use words as tokens. However, given a large text…