4 citations · 4 across the 2 of their papers we have counts for
1 paper · 1 filter
Ruihang Li, Yixuan Wei, Miaosen Zhang +3
High-quality data is crucial for the pre-training performance of large language models. Unfortunately, existing quality filtering methods rely on a known high-quality dataset as re…