Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models
Liangdong Wang, Bo-Wen Zhang, Chengwei Wu +7
We present CCI3.0-HQ (https://huggingface.co/datasets/BAAI/CCI3-HQ), a high-quality 500GB subset of the Chinese Corpora Internet 3.0 (CCI3.0)(https://huggingface.co/datasets/BAAI/C…
cs.CL2024
Beyond IID: Optimizing Instruction Learning from the Perspective of Instruction Interaction and Dependency
Hanyu Zhao, Li Du, Yiming Ju +2
With the availability of various instruction datasets, a pivotal challenge is how to effectively select and integrate these instructions to fine-tune large language models (LLMs).…