25 citations · 65 across the 14 of their papers we have counts for
18 papers
Analyzing Dialectical Biases in LLMs for Knowledge and Reasoning Benchmarks
Eileen Pan, Anna Seo Gyeong Choi, Maartje ter Hoeve +2
Large language models (LLMs) are ubiquitous in modern day natural language processing. However, previous work has shown degraded LLM performance for under-represented English diale…
Assessing the Role of Data Quality in Training Bilingual Language Models
Skyler Seto, Maartje ter Hoeve, Maureen de Seyssel +1
Bilingual and multilingual language models offer a promising path toward scaling NLP systems across diverse languages and users. However, their performance often varies wildly betw…
Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasks
Maureen de Seyssel, Jie Chi, Skyler Seto +3
We introduce a set of training-free ABX-style discrimination tasks to evaluate how multilingual language models represent language identity (form) and semantic content (meaning). I…
Analyzing the Effect of Linguistic Similarity on Cross-Lingual Transfer: Tasks and Experimental Setups Matter
Verena Blaschke, Masha Fedzechkina, Maartje ter Hoeve
Cross-lingual transfer is a popular approach to increase the amount of training data for NLP tasks in a low-resource context. However, the best strategy to decide which cross-lingu…
On the Way to LLM Personalization: Learning to Remember User Conversations
Lucie Charlotte Magister, Katherine Metcalf, Yizhe Zhang +1
Large Language Models (LLMs) have quickly become an invaluable assistant for a variety of tasks. However, their effectiveness is constrained by their ability to tailor responses to…
Training Bilingual LMs with Data Constraints in the Targeted Language
Skyler Seto, Maartje ter Hoeve, Richard He Bai +2
Large language models are trained on massive scrapes of the web, as required by current scaling laws. Most progress is made for English, given its abundance of high-quality pretrai…