activity
20242026
collaborators

6 papers

cs.CL2026

WhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Pretrained Speech Recognition Transformers

Akshat Pandey, Karun Kumar, Raphael Tang

Pretrained automatic speech recognition (ASR) models such as Whisper perform well but still need domain adaptation to handle unseen parlance. In many real-world settings, collectin…

cs.CL2026

DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining

Yutong Yan, Raphael Tang, Zhenyu Gao +2

Large language models pretrained on internet-scale data risk lookahead bias in forecasting tasks, as they may have already seen the true outcome during training. To address this, w…

cs.CL2026

The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining

Jiandong Shao, Raphael Tang, Crystina Zhang +4

Multilingual large language models achieve impressive cross-lingual performance despite largely monolingual pretraining. While bilingual data in pretraining corpora is widely belie…

cs.CL2025

Drawing Conclusions from Draws: Rethinking Preference Semantics in Arena-Style LLM Evaluation

Raphael Tang, Crystina Zhang, Wenyan Li +3

In arena-style evaluation of large language models (LLMs), two LLMs respond to a user query, and the user chooses the winning response or deems the "battle" a draw, resulting in an…

cs.CL2025

Multilingual Language Model Pretraining using Machine-translated Data

Jiayi Wang, Yao Lu, Maurice Weber +5

High-resource languages such as English, enables the pretraining of high-quality large language models (LLMs). The same can not be said for most other languages as LLMs still under…

cs.CL2024

Multilingual Pretraining Using a Large Corpus Machine-Translated from a Single Source Language

Jiayi Wang, Yao Lu, Maurice Weber +4

English, as a very high-resource language, enables the pretraining of high-quality large language models (LLMs). The same cannot be said for most other languages, as leading LLMs s…