6 papers · 1 filter
Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing
Filip Trhlik, Andrew Caines, Paula Buttery
Pre-trained language models (LMs) have, over the last few years, grown substantially in both societal adoption and training costs. This rapid growth in size has constrained progres…
Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction
Suchir Salhan, Hongyi Gu, Donya Rooein +5
Multi-turn dialogues between a child and a caregiver are characterized by a property called contingency - that is, prompt, direct, and meaningful exchanges between interlocutors. W…
BLiSS 1.0: Evaluating Bilingual Learner Competence in Second Language Small Language Models
Yuan Gao, Suchir Salhan, Andrew Caines +2
To bridge the gap between performance-oriented benchmarks and the evaluation of cognitively inspired models, we introduce BLiSS 1.0, a Benchmark of Learner Interlingual Syntactic S…
Prompting open-source and commercial language models for grammatical error correction of English learner text
Christopher Davis, Andrew Caines, Ãistein Andersen +6
Thanks to recent advances in generative AI, we are able to prompt large language models (LLMs) to produce texts which are fluent and grammatical. In addition, it has been shown tha…
From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes
Zébulon Goriely, Richard Diehl Martinez, Andrew Caines +2
Language models are typically trained on large corpora of text in their default orthographic form. However, this is not the only option; representing data as streams of phonemes ca…
Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing
Richard Diehl Martinez, Zebulon Goriely, Andrew Caines +2
Language models strongly rely on frequency information because they maximize the likelihood of tokens during pre-training. As a consequence, language models tend to not generalize…