papers

Publications (9)

cs.CL2026

Data Mixing for Large Language Models Pretraining: A Survey and Outlook

Zhuo Chen, Yuxuan Miao, Supryadi +1

Large language models (LLMs) rely on pretraining on massive and heterogeneous corpora, where training data composition has a decisive impact on training efficiency and downstream g…

cs.CL2026

LKValues: Aligning Large Language Models with Sri Lankan Societal Values

Nethmi Muthugala, Supryadi, Surangika Ranathunga +7

Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societi…

cs.CL2024

An Empirical Study on the Robustness of Massively Multilingual Neural Machine Translation

Supryadi, Leiyu Pan, Deyi Xiong

Massively multilingual neural machine translation (MMNMT) has been proven to enhance the translation quality of low-resource languages. In this paper, we empirically investigate th…

cs.CL2024

FuxiTranyu: A Multilingual Large Language Model Trained with Balanced Data

Haoran Sun, Renren Jin, Shaoyang Xu +10

Large language models (LLMs) have demonstrated prowess in a wide range of tasks. However, many LLMs exhibit significant performance discrepancies between high- and low-resource lan…

cs.CL2026

Pancasila-Dilemmas: Evaluating Large Language Models on Indonesian Human Value Dilemmas Grounded in Pancasila

Supryadi, Irfan, Julianti +4

The value alignment of large language models (LLMs) is crucial for ensuring responses align with human intention and value preferences. However, most evaluations of value alignment…

cs.AI2023

Is Robustness Transferable across Languages in Multilingual Neural Machine Translation?

Leiyu Pan, Supryadi, Deyi Xiong

Robustness, the ability of models to maintain performance in the face of perturbations, is critical for developing reliable NLP systems. Recent studies have shown promising results…

cs.CL2023

Evaluating Large Language Models: A Comprehensive Survey

Zishan Guo, Renren Jin, Chuang Liu +8

Large language models (LLMs) have demonstrated remarkable capabilities across a broad spectrum of tasks. They have attracted significant attention and been deployed in numerous dow…

cs.CV2025

Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia

Samuel Cahyawijaya, Holy Lovenia, Joel Ruben Antony Moniz +89

Southeast Asia (SEA) is a region of extraordinary linguistic and cultural diversity, yet it remains significantly underrepresented in vision-language (VL) research. This often resu…

cs.CL2024

Multilingual Large Language Models: A Systematic Survey

Shaolin Zhu, Supryadi, Shaoyang Xu +7

This paper provides a comprehensive survey of the latest research on multilingual large language models (MLLMs). MLLMs not only are able to understand and generate language across…