5 papers
Language Arithmetics: Towards Systematic Language Neuron Identification and Manipulation
Daniil Gurgurov, Katharina Trinley, Yusser Al Ghussin +3
Large language models (LLMs) exhibit strong multilingual abilities, yet the neural mechanisms behind language-specific processing remain unclear. We analyze language-specific neuro…
Modular Arithmetic: Language Models Solve Math Digit by Digit
Tanja Baeumel, Daniil Gurgurov, Yusser al Ghussin +2
While recent work has begun to uncover the internal strategies that Large Language Models (LLMs) employ for simple arithmetic tasks, a unified understanding of their underlying mec…
What Language(s) Does Aya-23 Think In? How Multilinguality Affects Internal Language Representations
Katharina Trinley, Toshiki Nakai, Tatiana Anikina +1
Large language models (LLMs) excel at multilingual tasks, yet their internal language processing remains poorly understood. We analyze how Aya-23-8B, a decoder-only LLM trained on…
Multilingual Large Language Models and Curse of Multilinguality
Daniil Gurgurov, Tanja Bäumel, Tatiana Anikina
Multilingual Large Language Models (LLMs) have gained large popularity among Natural Language Processing (NLP) researchers and practitioners. These models, trained on huge datasets…
The Lookahead Limitation: Why Multi-Operand Addition is Hard for LLMs
Tanja Baeumel, Josef van Genabith, Simon Ostermann
Autoregressive large language models (LLMs) exhibit impressive performance across various tasks but struggle with simple arithmetic, such as addition of two or more operands. We sh…