23 citations · 24 across the 3 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2022★ 1 cited
Neural Token Segmentation for High Token-Internal Complexity
Idan Brusilovsky, Reut Tsarfaty
Tokenizing raw texts into word units is an essential pre-processing step for critical tasks in the NLP pipeline such as tagging, parsing, named entity recognition, and more. For mo…
cs.CL2021★ 23 cited
AlephBERT:A Hebrew Large Pre-Trained Language Model to Start-off your Hebrew NLP Application With
Amit Seker, Elron Bandel, Dan Bareket +3
Large Pre-trained Language Models (PLMs) have become ubiquitous in the development of language understanding technology and lie at the heart of many artificial intelligence advance…