Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
StochasTok: Improving Fine-Grained Subword Understanding in LLMs
Anya Sims, Thom Foster, Klara Kaleb +5
Subword-level understanding is integral to numerous tasks, including understanding multi-digit numbers, spelling mistakes, abbreviations, rhyming, and wordplay. Despite this, curre…
cs.CL2026
Stochasticity in Tokenisation Improves Robustness
Sophie Steger, Rui Li, Sofiane Ennadir +4
The widespread adoption of large language models (LLMs) has increased concerns about their robustness. Vulnerabilities in perturbations of tokenisation of the input indicate that m…