5 papers · 1 filter
A Large-Scale Chinese Knowledge Graph-Text Alignment Dataset for Benchmarking Knowledge-Grounded LLMs
Chengwei Wu, Jiapu Wang, Mingyang Gao +10
Reliable evaluation of knowledge-grounded Large Language Models (LLMs) in Chinese requires resources that explicitly align Chinese-language text with verifiable Knowledge Graph (KG…
CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models
Liangdong Wang, Bo-Wen Zhang, Chengwei Wu +7
We present CCI3.0-HQ (https://huggingface.co/datasets/BAAI/CCI3-HQ), a high-quality 500GB subset of the Chinese Corpora Internet 3.0 (CCI3.0)(https://huggingface.co/datasets/BAAI/C…
Beyond IID: Optimizing Instruction Learning from the Perspective of Instruction Interaction and Dependency
Hanyu Zhao, Li Du, Yiming Ju +2
With the availability of various instruction datasets, a pivotal challenge is how to effectively select and integrate these instructions to fine-tune large language models (LLMs).…
AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies
Bo-Wen Zhang, Liangdong Wang, Ye Yuan +24
In recent years, with the rapid application of large language models across various fields, the scale of these models has gradually increased, and the resources required for their…
Harnessing Multi-Role Capabilities of Large Language Models for Open-Domain Question Answering
Hongda Sun, Yuxuan Liu, Chengwei Wu +5
Open-domain question answering (ODQA) has emerged as a pivotal research spotlight in information systems. Existing methods follow two main paradigms to collect evidence: (1) The \t…