1 paper
Zheheng Luo, Xin Zhang, Xiao Liu +4
It is well-known that a diverse corpus is critical for training large language models, which are typically constructed from a mixture of various domains. In general, previous effor…