Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures
Ziyun Qiao, Yue Min, Ruining Chen +1
Most data-mixing methods assume the corpus has already been partitioned into groups, and the choice of those groups determines what a mixer can express. Existing labels, including…
cs.LG2026
GRASP: Geometry-aware Residual Alignment for Scalable Pretraining Data Attribution
Yue Min, Ruining Chen, Yujun Li
Scalable data attribution methods typically assign isolated utility scores to individual training examples. This prevalent additive assumption fundamentally fails to capture critic…
cs.LG2026
GEM: Geometric Entropy Mixing for Optimal LLM Data Curation
Yue Min, Ziyun Qiao, Ruining Chen +1
LLM pre-training efficacy increasingly depends on data composition rather than sheer volume. Yet, optimal mixing is hindered by categorization flaws: human taxonomies suffer from o…