3 papers
cs.CL2025
Once Upon a Time: Interactive Learning for Storytelling with Small Language Models
Jonas Mayer Martins, Ali Hamza Bashir, Muhammad Rehan Khalid +1
Children efficiently acquire language not just by listening, but by interacting with others in their social environment. Conversely, large language models are typically trained wit…
cs.CL2024
From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes
Zébulon Goriely, Richard Diehl Martinez, Andrew Caines +2
Language models are typically trained on large corpora of text in their default orthographic form. However, this is not the only option; representing data as streams of phonemes ca…
cs.CL2024
Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing
Richard Diehl Martinez, Zebulon Goriely, Andrew Caines +2
Language models strongly rely on frequency information because they maximize the likelihood of tokens during pre-training. As a consequence, language models tend to not generalize…