6 papers
WhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Pretrained Speech Recognition Transformers
Akshat Pandey, Karun Kumar, Raphael Tang
Pretrained automatic speech recognition (ASR) models such as Whisper perform well but still need domain adaptation to handle unseen parlance. In many real-world settings, collectin…
DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining
Yutong Yan, Raphael Tang, Zhenyu Gao +2
Large language models pretrained on internet-scale data risk lookahead bias in forecasting tasks, as they may have already seen the true outcome during training. To address this, w…
The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining
Jiandong Shao, Raphael Tang, Crystina Zhang +4
Multilingual large language models achieve impressive cross-lingual performance despite largely monolingual pretraining. While bilingual data in pretraining corpora is widely belie…
Drawing Conclusions from Draws: Rethinking Preference Semantics in Arena-Style LLM Evaluation
Raphael Tang, Crystina Zhang, Wenyan Li +3
In arena-style evaluation of large language models (LLMs), two LLMs respond to a user query, and the user chooses the winning response or deems the "battle" a draw, resulting in an…
Multilingual Language Model Pretraining using Machine-translated Data
Jiayi Wang, Yao Lu, Maurice Weber +5
High-resource languages such as English, enables the pretraining of high-quality large language models (LLMs). The same can not be said for most other languages as LLMs still under…
Multilingual Pretraining Using a Large Corpus Machine-Translated from a Single Source Language
Jiayi Wang, Yao Lu, Maurice Weber +4
English, as a very high-resource language, enables the pretraining of high-quality large language models (LLMs). The same cannot be said for most other languages, as leading LLMs s…