3 papers
cs.AI2024
Harnessing Diversity for Important Data Selection in Pretraining Large Language Models
Chi Zhang, Huaping Zhong, Kuan Zhang +10
Data selection is of great significance in pre-training large language models, given the variation in quality within the large-scale available training corpora. To achieve this, re…
cs.CL2024
InternLM2 Technical Report
Zheng Cai, Maosong Cao, Haojiong Chen +97
The evolution of Large Language Models (LLMs) like ChatGPT and GPT-4 has sparked discussions on the advent of Artificial General Intelligence (AGI). However, replicating such advan…
cs.CL2024
WanJuan-CC: A Safe and High-Quality Open-sourced English Webtext Dataset
Jiantao Qiu, Haijun Lv, Zhenjiang Jin +23
This paper presents WanJuan-CC, a safe and high-quality open-sourced English webtext dataset derived from Common Crawl data. The study addresses the challenges of constructing larg…