2 citations · 2 across the 3 of their papers we have counts for
3 papers
BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data
Jaap Jumelet, Abdellah Fourtassi, Akari Haga +23
We present BabyBabelLM, a multilingual collection of datasets modeling the language a person observes from birth until they acquire a native language. We curate developmentally pla…
BERTtime Stories: Investigating the Role of Synthetic Story Data in Language Pre-training
Nikitas Theodoropoulos, Giorgos Filandrianos, Vassilis Lyberatos +2
We describe our contribution to the Strict and Strict-Small tracks of the 2nd iteration of the BabyLM Challenge. The shared task is centered around efficient pre-training given dat…
From {Solution Synthesis} to {Student Attempt Synthesis} for Block-Based Visual Programming Tasks
Adish Singla, Nikitas Theodoropoulos
Block-based visual programming environments are increasingly used to introduce computing concepts to beginners. Given that programming tasks are open-ended and conceptual, novice s…