7 citations · 7 across the 1 of their papers we have counts for
2 papers
cs.CL2022★ 7 cited
Breaking Character: Are Subwords Good Enough for MRLs After All?
Omri Keren, Tal Avinari, Reut Tsarfaty +1
Large pretrained language models (PLMs) typically tokenize the input string into contiguous subwords before any pretraining or inference. However, previous studies have claimed tha…
cs.CL2021
ParaShoot: A Hebrew Question Answering Dataset
Omri Keren, Omer Levy
NLP research in Hebrew has largely focused on morphology and syntax, where rich annotated datasets in the spirit of Universal Dependencies are available. Semantic datasets, however…