2 papers
cs.CL2026
Joint Optimization for Greedy Longest-match Tokenization
Adhiraj Singh, Deepanshu Mody, Ghina Al Shdaifat +4
Recent work has shown that subword vocabularies can be trained to optimize compression for a specific inference rule rather than relying on greedy heuristics such as Byte Pair Enco…
cs.CL2026
Tokenization with Split Trees
Craig W. Schmidt, Michael Krumdick, Adam Wiemerslage +4
We introduce Tokenization with Split Trees (ToaST), a subword tokenization method that directly optimizes compression under a new recursive inference procedure. ToaST greedily spli…