2 papers
cs.CL2026
Reducing Tokenization Premiums for Low-Resource Languages
Geoffrey Churchill, Steven Skiena
Relative to English, low-resource languages suffer from substantial tokenization premiums in modern LMs, meaning that it generally requires several times as many tokens to encode a…
cs.CL2025
Word Definitions from Large Language Models
Bach Pham, JuiHsuan Wong, Samuel Kim +2
Dictionary definitions are historically the arbitrator of what words mean, but this primacy has come under threat by recent progress in NLP, including word embeddings and generativ…