2 papers
cs.CL2025
Learning from the Best, Differently: A Diversity-Driven Rethinking on Data Selection
Hongyi He, Xiao Liu, Zhenghao Lin +6
High-quality pre-training data is crutial for large language models, where quality captures factual reliability and semantic value, and diversity ensures broad coverage and distrib…
cs.CL2025
Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models
Zhenghao Lin, Zihao Tang, Xiao Liu +31
We introduce Sigma, an efficient large language model specialized for the system domain, empowered by a novel architecture including DiffQKV attention, and pre-trained on our metic…