From the 1 of 5 linked papers with an AI index.
1 paper · 1 filter
Zhili Feng, Tanya Marwah, Nicolo Fusi +2
Modern large language models use a fixed tokenizer to effectively compress text drawn from a source domain. However, applying the same tokenizer to a new target domain often leads…