3 citations · 3 across the 1 of their papers we have counts for
3 papers
cs.CL2024
Upsample or Upweight? Balanced Training on Heavily Imbalanced Datasets
Tianjian Li, Haoran Xu, Weiting Tan +2
Data abundance across different domains exhibits a long-tailed distribution: few domains have abundant data, while most face data scarcity. Our work focuses on a multilingual setti…
cs.CL2024
Streaming Sequence Transduction through Dynamic Compression
Weiting Tan, Yunmo Chen, Tongfei Chen +5
We introduce STAR (Stream Transduction with Anchor Representations), a novel Transformer-based model designed for efficient sequence-to-sequence transduction over streams. STAR dyn…
cs.CL2024★ 3 cited
The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts
Lingfeng Shen, Weiting Tan, Sihao Chen +6
As the influence of large language models (LLMs) spans across global communities, their safety challenges in multilingual settings become paramount for alignment research. This pap…