5 papers
PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
Daryna Dementieva, Nikolay Babakov, Kathy Hämmerl +14
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resourc…
Mask and You Shall Receive: Optimizing Masked Language Modeling For Pretraining BabyLMs
Lukas Edman, Alexander Fraser
We describe our strategy for the 2025 edition of the BabyLM Challenge. Our main contribution is that of an improved form of Masked Language Modeling (MLM), which adapts the probabi…
EXECUTE: A Multilingual Benchmark for LLM Token Understanding
Lukas Edman, Helmut Schmid, Alexander Fraser
The CUTE benchmark showed that LLMs struggle with character understanding in English. We extend it to more languages with diverse scripts and writing systems, introducing EXECUTE.…
Are BabyLMs Second Language Learners?
Lukas Edman, Lisa Bylinina, Faeze Ghorbanpour +1
This paper describes a linguistically-motivated approach to the 2024 edition of the BabyLM Challenge (Warstadt et al. 2023). Rather than pursuing a first language learning (L1) par…
CUTE: Measuring LLMs' Understanding of Their Tokens
Lukas Edman, Helmut Schmid, Alexander Fraser
Large Language Models (LLMs) show remarkable performance on a wide variety of tasks. Most LLMs split text into multi-character tokens and process them as atomic units without direc…