Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles
Buu Phan, Brandon Amos, Itai Gat +3
Tokenization is associated with many poorly understood shortcomings in language models (LMs), yet remains an important component for long sequence scaling purposes. This work studi…
cs.CL2024
Understanding and Mitigating Tokenization Bias in Language Models
Buu Phan, Marton Havasi, Matthew Muckley +1
State-of-the-art language models are autoregressive and operate on subword units known as tokens. Specifically, one must encode the conditioning string into a list of tokens before…