3 citations · 3 across the 3 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values
P Team, Siwei Wu, Jincheng Ren +29
Aligning large language models (LLMs) with human preferences has achieved remarkable success. However, existing Chinese preference datasets are limited by small scale, narrow domai…
cs.CL2024★ 3 cited
D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
Haoran Que, Jiaheng Liu, Ge Zhang +13
Continual Pre-Training (CPT) on Large Language Models (LLMs) has been widely used to expand the model's fundamental understanding of specific downstream domains (e.g., math and cod…
cs.CL2024
E^2-LLM: Efficient and Extreme Length Extension of Large Language Models
Jiaheng Liu, Zhiqi Bai, Yuanxing Zhang +11
Typically, training LLMs with long context sizes is computationally expensive, requiring extensive training hours and GPU resources. Existing long-context extension methods usually…