Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation
Buu Phan, Ashish Khisti, Karen Ullrich
Computing next-token likelihood ratios between two language models (LMs) is a standard task in training paradigms such as knowledge distillation. Since this requires both models to…
cs.CL2025
Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles
Buu Phan, Brandon Amos, Itai Gat +3
Tokenization is associated with many poorly understood shortcomings in language models (LMs), yet remains an important component for long sequence scaling purposes. This work studi…