1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2024
Understanding and Mitigating Tokenization Bias in Language Models
Buu Phan, Marton Havasi, Matthew Muckley +1
State-of-the-art language models are autoregressive and operate on subword units known as tokens. Specifically, one must encode the conditioning string into a list of tokens before…
cs.IT2024★ 1 cited
Importance Matching Lemma for Lossy Compression with Side Information
Buu Phan, Ashish Khisti, Christos Louizos
We propose two extensions to existing importance sampling based methods for lossy compression. First, we introduce an importance sampling based compression scheme that is a variant…