3 papers
cs.CL2024
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
Pavel Chizhov, Catherine Arnett, Elizaveta Korotkova +1
Language models can largely benefit from efficient tokenization. However, they still mostly utilize the classical BPE algorithm, a simple and reliable method. This has been shown t…
cs.CL2022
Pragmatic Constraint on Distributional Semantics
Elizaveta Zhemchuzhina, Nikolai Filippov, Ivan P. Yamshchikov
This paper studies the limits of language models' statistical learning in the context of Zipf's law. First, we demonstrate that Zipf-law token distribution emerges irrespective of…
cs.CL2022
Moving Other Way: Exploring Word Mover Distance Extensions
Ilya Smirnov, Ivan P. Yamshchikov
The word mover's distance (WMD) is a popular semantic similarity metric for two texts. This position paper studies several possible extensions of WMD. We experiment with the freque…