1 paper
Lin Zheng, Xinyu Li, Qian Liu +2
Modern language models are trained almost exclusively on token sequences produced by a fixed tokenizer, an external lossless compressor often over UTF-8 byte sequences, thereby cou…