collaborators

9 papers

cs.CL2026

Disentangling Linguistic Relatedness from Task Alignment in Cross-Lingual Transfer

Ahmed Haj Ahmed, Ruochen Zhang, Alvin Grissom

We study cross-lingual transfer by fine-tuning seven large language models (4B--671B parameters) on Arabic and evaluating zero-shot reading comprehension on Semitic languages and n…

cs.AI2026

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model

Samuel Cahyawijaya, Peerat Limkonchotiwat, Tack Hwa Wong +45

While the field of vision-language (VL) has achieved remarkable success in integrating visual and textual information across multiple languages and domains, there is still no dedic…

cs.CL2026

Can Large Language Models Understand, Reason About, and Generate Code-Switched Text?

Genta Indra Winata, David Anugraha, Patrick Amadeus Irawan +15

Code-switching is a pervasive phenomenon in multilingual communication, yet the robustness of large language models (LLMs) in mixed-language settings remains insufficiently underst…

cs.LG2025

Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability

Genta Indra Winata, David Anugraha, Emmy Liu +17

High-quality datasets are fundamental to training and evaluating machine learning models, yet their creation-especially with accurate human annotations-remains a significant challe…

cs.CL2025

Paths Not Taken: Understanding and Mending the Multilingual Factual Recall Pipeline

Meng Lu, Ruochen Zhang, Carsten Eickhoff +1

Multilingual large language models (LLMs) often exhibit factual inconsistencies across languages, with significantly better performance in factual recall tasks in English than in o…

cs.CL2025

Crosslingual Reasoning through Test-Time Scaling

Zheng-Xin Yong, M. Farid Adilazuarda, Jonibek Mansurov +7

Reasoning capabilities of large language models are primarily studied for English, even when pretrained models are multilingual. In this work, we investigate to what extent English…