39 citations · 126 across the 13 of their papers we have counts for
12 papers · 1 filter
From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora
Yingli Shen, Wen Lai, Shuo Wang +4
Continued pretraining and instruction tuning on large-scale multilingual data have proven to be effective in scaling large language models (LLMs) to low-resource languages. However…
Mastering Text, Code and Math Simultaneously via Fusing Highly Specialized Language Models
Ning Ding, Yulin Chen, Ganqu Cui +6
Underlying data distributions of natural language, programming code, and mathematical symbols vary vastly, presenting a complex challenge for large language models (LLMs) that stri…
UltraLink: An Open-Source Knowledge-Enhanced Multilingual Supervised Fine-tuning Dataset
Haoyu Wang, Shuo Wang, Yukun Yan +9
Open-source large language models (LLMs) have gained significant strength across diverse fields. Nevertheless, the majority of studies primarily concentrate on English, with only l…
READIN: A Chinese Multi-Task Benchmark with Realistic and Diverse Input Noises
Chenglei Si, Zhengyan Zhang, Yingfa Chen +3
For many real-world applications, the user-generated inputs usually contain various noises due to speech recognition errors caused by linguistic variations1 or typographical errors…
Fuse It More Deeply! A Variational Transformer with Layer-Wise Latent Variable Inference for Text Generation
Jinyi Hu, Xiaoyuan Yi, Wenhao Li +2
The past several years have witnessed Variational Auto-Encoder's superiority in various text generation tasks. However, due to the sequential nature of the text, auto-regressive de…
YACLC: A Chinese Learner Corpus with Multidimensional Annotation
Yingying Wang, Cunliang Kong, Liner Yang +8
Learner corpus collects language data produced by L2 learners, that is second or foreign-language learners. This resource is of great relevance for second language acquisition rese…