10 papers
OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents
Andrea Caciolai, Pere-Lluís Huguet Cabot, Chierh Cheng +11
Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI…
Extracting Small Translation Specialists from LLMs by Aggressively Pruning Experts
Liu O. Martin, Lucas Bandarkar, Nanyun Peng
Modern large language models (LLMs) achieve state-of-the-art machine translation performance, but they do so as broad generalists largely trained for many tasks and capabilities un…
Knowledge Localization in Mixture-of-Experts LLMs Using Cross-Lingual Inconsistency
Lucas Bandarkar, Alan Ansell, Trevor Cohn
Modern LLMs continue to exhibit significant variance in behavior across languages, such as being able to recall factual information in some languages but not others. While typicall…
Large Reasoning Models Struggle to Transfer Parametric Knowledge Across Scripts
Lucas Bandarkar, Alan Ansell, Trevor Cohn
In this work, we analyze shortcomings in cross-lingual knowledge transfer in large, modern reasoning LLMs. We demonstrate that the perceived gap in knowledge transfer is primarily…
Multilingual Routing in Mixture-of-Experts
Lucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz +2
Mixture-of-Experts (MoE) architectures have become the key to scaling modern LLMs, yet little is understood about how their sparse routing dynamics respond to multilingual data. In…
Translation as a Scalable Proxy for Multilingual Evaluation
Sheriff Issaka, Erick Rosas Gonzalez, Lieqi Liu +6
The rapid proliferation of LLMs has created a critical evaluation paradox: while LLMs claim multilingual proficiency, comprehensive non-machine-translated benchmarks exist for fewe…