1 citations · 2 across the 9 of their papers we have counts for
5 papers · 1 filter
Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization
Qian Kou, Xiaofeng Shi, Xiaosong Qiu +1
Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as doc…
Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM Accuracy
Xiaofeng Shi, Qian Kou, Yuduo Li +1
With the rapid advancement of Large Language Models (LLMs), the Chain-of-Thought (CoT) component has become significant for complex reasoning tasks. However, in conventional Superv…
CareBot: A Pioneering Full-Process Open-Source Medical Language Model
Lulu Zhao, Weihao Zeng, Xiaofeng Shi +1
Recently, both closed-source LLMs and open-source communities have made significant strides, outperforming humans in various general domains. However, their performance in specific…
CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models
Liangdong Wang, Bo-Wen Zhang, Chengwei Wu +7
We present CCI3.0-HQ (https://huggingface.co/datasets/BAAI/CCI3-HQ), a high-quality 500GB subset of the Chinese Corpora Internet 3.0 (CCI3.0)(https://huggingface.co/datasets/BAAI/C…
Aqulia-Med LLM: Pioneering Full-Process Open-Source Medical Language Models
Lulu Zhao, Weihao Zeng, Xiaofeng Shi +3
Recently, both closed-source LLMs and open-source communities have made significant strides, outperforming humans in various general domains. However, their performance in specific…