6 citations · 7 across the 15 of their papers we have counts for
5 papers · 1 filter
FastMix: Fast Data Mixture Optimization via Gradient Descent
Haoru Tan, Sitong Wu, Yanfeng Chen +5
While large and diverse datasets have driven recent advances in large models, identifying the optimal data mixture for pre-training and post-training remains a significant open pro…
MC#: Mixture Compressor for Mixture-of-Experts Large Models
Wei Huang, Yue Liao, Yukang Chen +6
Mixture-of-Experts (MoE) effectively scales large language models (LLMs) and vision-language models (VLMs) by increasing capacity through sparse activation. However, preloading all…
Understanding Data Influence with Differential Approximation
Haoru Tan, Sitong Wu, Xiuzhe Wu +5
Data plays a pivotal role in the groundbreaking advancements in artificial intelligence. The quantitative analysis of data significantly contributes to model training, enhancing bo…
Mixture Compressor for Mixture-of-Experts LLMs Gains More
Wei Huang, Yue Liao, Jianhui Liu +6
Mixture-of-Experts large language models (MoE-LLMs) marks a significant step forward of language models, however, they encounter two critical challenges in practice: 1) expert para…
Data Pruning via Moving-one-Sample-out
Haoru Tan, Sitong Wu, Fei Du +4
In this paper, we propose a novel data-pruning approach called moving-one-sample-out (MoSo), which aims to identify and remove the least informative samples from the training set.…