1 citations · 1 across the 3 of their papers we have counts for
3 papers
BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization
Sander Land, Catherine Arnett
Byte Pair Encoding (BPE) tokenizers, widely used in Large Language Models, face challenges in multilingual settings, including penalization of non-Western scripts and the creation…
Toxicity of the Commons: Curating Open-Source Pre-Training Data
Catherine Arnett, Eliot Jones, Ivan P. Yamshchikov +1
Open-source large language models are becoming increasingly available and popular among researchers and practitioners. While significant progress has been made on open-weight model…
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
Pavel Chizhov, Catherine Arnett, Elizaveta Korotkova +1
Language models can largely benefit from efficient tokenization. However, they still mostly utilize the classical BPE algorithm, a simple and reliable method. This has been shown t…