17 papers
Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining
Felicia Körner, Maria Matveev, Florian Eichin +3
Large language models exhibit impressive cross-lingual capabilities. However, prior work analyzes this phenomenon through isolated factors and at sparse points during training, lim…
AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages
Hao Yu, Tianyi Xu, Michael A. Hedderich +3
Large language models (LLMs) are increasingly multilingual, yet open models continue to underperform relative to proprietary systems, with the gap most pronounced for African langu…
ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior
Florian Eichin, Yupei Du, Philipp Mondorf +3
Post-hoc interpretability methods typically attribute a model's behavior to its components, data, or training trajectory in isolation, and are often tied to a particular level of g…
Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
Raoyuan Zhao, Yihong Liu, Lena Altinger +2
Large language models (LLMs) are increasingly deployed in multilingual, real-world applications with user inputs -- naturally introducing \emph{typographical errors} (typos). Yet m…
Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners
Yihong Liu, Raoyuan Zhao, Hinrich Schütze +1
Large reasoning models (LRMs) achieve strong performance on mathematical reasoning tasks, often attributed to their capability to generate explicit chain-of-thought (CoT) explanati…
ToMigo: Interpretable Design Concept Graphs for Aligning Generative AI with Creative Intent
Lena Hegemann, Xinyi Wen, Michael A. Hedderich +2
Generative AI often produces results misaligned with user intentions, for example, resolving ambiguous prompts in unexpected ways. Despite existing approaches to clarify intent, a…