collaborators

6 papers

cs.CL2025

Understanding Subword Compositionality of Large Language Models

Qiwei Peng, Yekun Chai, Anders Søgaard

Large language models (LLMs) take sequences of subwords as input, requiring them to effective compose subword representations into meaningful word-level representations. In this pa…

cs.CL2025

Debiasing Multilingual LLMs in Cross-lingual Latent Space

Qiwei Peng, Guimin Hu, Yekun Chai +1

Debiasing techniques such as SentDebias aim to reduce bias in large language models (LLMs). Previous studies have evaluated their cross-lingual transferability by directly applying…

cs.CL2025

SemEval-2025 Task 7: Multilingual and Crosslingual Fact-Checked Claim Retrieval

Qiwei Peng, Robert Moro, Michal Gregor +7

The rapid spread of online disinformation presents a global challenge, and machine learning has been widely explored as a potential solution. However, multilingual settings and low…

cs.CL2024

Tokenization Falling Short: On Subword Robustness in Large Language Models

Yekun Chai, Yewei Fang, Qiwei Peng +1

Language models typically tokenize raw text into sequences of subword identifiers from a predefined vocabulary, a process inherently sensitive to typographical errors, length varia…

cs.CL2024

Concept Space Alignment in Multilingual LLMs

Qiwei Peng, Anders Søgaard

Multilingual large language models (LLMs) seem to generalize somewhat across languages. We hypothesize this is a result of implicit vector space alignment. Evaluating such alignmen…

cs.CL2024

FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food Culture

Wenyan Li, Xinyu Zhang, Jiaang Li +9

Food is a rich and varied dimension of cultural heritage, crucial to both individuals and social groups. To bridge the gap in the literature on the often-overlooked regional divers…