1 paper
Bar Gazit, Shaltiel Shmidman, Avi Shmidman +1
Common subword tokenization algorithms like BPE and UnigramLM assume that text can be split into meaningful units by concatenative measures alone. This is not true for languages su…