3 papers
cs.CL2025
ByteSpan: Information-Driven Subword Tokenisation
Zébulon Goriely, Suchir Salhan, Pietro Lesci +2
Recent dynamic tokenisation methods operate directly on bytes and pool their latent representations into patches. This bears similarities to computational models of word segmentati…
cs.CL2025
PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs
Oskar van der Wal, Pietro Lesci, Max Muller-Eberstein +4
The stability of language model pre-training and its effects on downstream performance are still understudied. Prior work shows that the training process can yield significantly di…
cs.CL2024
Tending Towards Stability: Convergence Challenges in Small Language Models
Richard Diehl Martinez, Pietro Lesci, Paula Buttery
Increasing the number of parameters in language models is a common strategy to enhance their performance. However, smaller language models remain valuable due to their lower operat…