2 papers
cs.LG2026
HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures
Ziyun Qiao, Yue Min, Ruining Chen +1
Most data-mixing methods assume the corpus has already been partitioned into groups, and the choice of those groups determines what a mixer can express. Existing labels, including…
cs.LG2026
GEM: Geometric Entropy Mixing for Optimal LLM Data Curation
Yue Min, Ziyun Qiao, Ruining Chen +1
LLM pre-training efficacy increasingly depends on data composition rather than sheer volume. Yet, optimal mixing is hindered by categorization flaws: human taxonomies suffer from o…