2 citations · 3 across the 2 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025★ 1 cited
Group-Level Data Selection for Efficient Pretraining
Zichun Yu, Fei Peng, Jie Lei +3
In this paper, we introduce Group-MATES, an efficient group-level data selection approach to optimize the speed-quality frontier of language model pretraining. Specifically, Group-…
cs.CL2022★ 2 cited
Automatic Label Sequence Generation for Prompting Sequence-to-sequence Models
Zichun Yu, Tianyu Gao, Zhengyan Zhang +4
Prompting, which casts downstream applications as language modeling tasks, has shown to be sample efficient compared to standard fine-tuning with pre-trained models. However, one p…