3 papers
cs.CL2026
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…
cs.CL2026
Positional Cognitive Specialization: Where Do LLMs Learn To Comprehend and Speak Your Language?
Luis Frentzen Salim, Lun-Wei Ku, Hsing-Kuo Kenneth Pao
Adapting large language models (LLMs) to new languages is an expensive and opaque process. Understanding how language models acquire new languages and multilingual abilities is key…
cs.CL2026
Beyond Many-Shot Translation: Scaling In-Context Demonstrations For Low-Resource Machine Translation
Luis Frentzen Salim, Esteban Carlin, Alexandre Morinvil +2
Building machine translation (MT) systems for low-resource languages is notably difficult due to the scarcity of high-quality data. Although Large Language Models (LLMs) have impro…