4 papers
BabyLM's First Words: Word Segmentation as a Phonological Probing Task
Zébulon Goriely, Paula Buttery
Language models provide a key framework for studying linguistic theories based on prediction, but phonological analysis using large language models (LLMs) is difficult; there are f…
IPA-CHILDES & G2P+: Feature-Rich Resources for Cross-Lingual Phonology and Phonemic Language Modeling
Zébulon Goriely, Paula Buttery
In this paper, we introduce two resources: (i) G2P+, a tool for converting orthographic datasets to a consistent phonemic representation; and (ii) IPA CHILDES, a phonemic dataset o…
From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes
Zébulon Goriely, Richard Diehl Martinez, Andrew Caines +2
Language models are typically trained on large corpora of text in their default orthographic form. However, this is not the only option; representing data as streams of phonemes ca…
CLIMB: Curriculum Learning for Infant-inspired Model Building
Richard Diehl Martinez, Zebulon Goriely, Hope McGovern +4
We describe our team's contribution to the STRICT-SMALL track of the BabyLM Challenge. The challenge requires training a language model from scratch using only a relatively small t…