activity
20242026
collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2026

CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks

Peiqin Lin, Chenyang Lyu, Wenjiang Luo +22

Large language models (LLMs) are now deployed worldwide, inspiring a surge of benchmarks that measure their multilingual and multicultural abilities. However, these benchmarks prio…

cs.CL2025

LangSAMP: Language-Script Aware Multilingual Pretraining

Yihong Liu, Haotian Ye, Chunlan Ma +2

Recent multilingual pretrained language models (mPLMs) often avoid using language embeddings -- learnable vectors assigned to individual languages. However, this places a significa…

cs.CL2024

How Transliterations Improve Crosslingual Alignment

Yihong Liu, Mingyang Wang, Amir Hossein Kargaran +6

Recent studies have shown that post-aligning multilingual pretrained language models (mPLMs) using alignment objectives on both original and transliterated data can improve crossli…

cs.CL2024

TransMI: A Framework to Create Strong Baselines from Multilingual Pretrained Language Models for Transliterated Data

Yihong Liu, Chunlan Ma, Haotian Ye +1

Transliterating related languages that use different scripts into a common script is effective for improving crosslingual transfer in downstream tasks. However, this methodology of…

cs.CL2024

Exploring the Role of Transliteration in In-Context Learning for Low-resource Languages Written in Non-Latin Scripts

Chunlan Ma, Yihong Liu, Haotian Ye +1

Decoder-only large language models (LLMs) excel in high-resource languages across various tasks through few-shot or even zero-shot in-context learning (ICL). However, their perform…

cs.CL2024

Taxi1500: A Multilingual Dataset for Text Classification in 1500 Languages

Chunlan Ma, Ayyoob ImaniGooghari, Haotian Ye +3

While natural language processing tools have been developed extensively for some of the world's languages, a significant portion of the world's over 7000 languages are still neglec…