Showing cs.CLShow all
3 papers · 1 filter
cs.CL2024
ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws
Ruihang Li, Yixuan Wei, Miaosen Zhang +3
High-quality data is crucial for the pre-training performance of large language models. Unfortunately, existing quality filtering methods rely on a known high-quality dataset as re…
cs.CL2024
Xwin-LM: Strong and Scalable Alignment Practice for LLMs
Bolin Ni, JingCheng Hu, Yixuan Wei +4
In this work, we present Xwin-LM, a comprehensive suite of alignment methodologies for large language models (LLMs). This suite encompasses several key techniques, including superv…
cs.CL2024
Common 7B Language Models Already Possess Strong Math Capabilities
Chen Li, Weiqi Wang, Jingcheng Hu +5
Mathematical capabilities were previously believed to emerge in common language models only at a very large scale or require extensive math-related pre-training. This paper shows t…