Showing cs.CLShow all
3 papers · 1 filter
cs.CL2024
From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes
Zébulon Goriely, Richard Diehl Martinez, Andrew Caines +2
Language models are typically trained on large corpora of text in their default orthographic form. However, this is not the only option; representing data as streams of phonemes ca…
cs.CL2023
CLIMB: Curriculum Learning for Infant-inspired Model Building
Richard Diehl Martinez, Zebulon Goriely, Hope McGovern +4
We describe our team's contribution to the STRICT-SMALL track of the BabyLM Challenge. The challenge requires training a language model from scratch using only a relatively small t…
cs.CL2023
Finding the Needle in a Haystack: Unsupervised Rationale Extraction from Long Text Classifiers
Kamil Bujel, Andrew Caines, Helen Yannakoudakis +1
Long-sequence transformers are designed to improve the representation of longer texts by language models and their performance on downstream document-level tasks. However, not much…