1 paper
Yuval Reif, Guy Kaplan, Roy Schwartz
Large language models (LLMs) often encode word-form variation (e.g., walk vs. walked) as linear directions in the embedding space. However, standard tokenization algorithms treat s…