activity
20242026
collaborators

9 papers

cs.CL2026

Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions

Anastasiia Sedova, Natalie Schluter, Skyler Seto +1

Cross-lingual knowledge transfer is critical for building high-performing multilingual language models for languages with insufficient training data. When target language data is s…

cs.CL2026

How Value Induction Reshapes LLM Behaviour

Arnav Arora, Natalie Schluter, Katherine Metcalf +1

Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and empathy, and values, such as h…

cs.CL2025

Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasks

Maureen de Seyssel, Jie Chi, Skyler Seto +3

We introduce a set of training-free ABX-style discrimination tasks to evaluate how multilingual language models represent language identity (form) and semantic content (meaning). I…

cs.CL2025

Analyzing Dialectical Biases in LLMs for Knowledge and Reasoning Benchmarks

Eileen Pan, Anna Seo Gyeong Choi, Maartje ter Hoeve +2

Large language models (LLMs) are ubiquitous in modern day natural language processing. However, previous work has shown degraded LLM performance for under-represented English diale…

cs.CL2025

Assessing the Role of Data Quality in Training Bilingual Language Models

Skyler Seto, Maartje ter Hoeve, Maureen de Seyssel +1

Bilingual and multilingual language models offer a promising path toward scaling NLP systems across diverse languages and users. However, their performance often varies wildly betw…

cs.CL2025

GrammaMT: Improving Machine Translation with Grammar-Informed In-Context Learning

Rita Ramos, Everlyn Asiko Chimoto, Maartje ter Hoeve +1

We introduce GrammaMT, a grammatically-aware prompting approach for machine translation that uses Interlinear Glossed Text (IGT), a common form of linguistic description providing…