4 papers
Range Penalization: Theoretical Insights with Applications in Federated Learning
Yiyuan She, Zhaojun Hu, Yifan Sun
This paper introduces range regularization for federated learning with linear systematic components to enhance statistical accuracy and induce cross-client regularity conducive to…
HMS-BERT: Hybrid Multi-Task Self-Training for Multilingual and Multi-Label Cyberbullying Detection
Zixin Feng, Xinying Cui, Yifan Sun +5
Cyberbullying on social media is inherently multilingual and multi-faceted, where abusive behaviors often overlap across multiple categories. Existing methods are commonly limited…
Influence-Preserving Proxies for Gradient-Based Data Selection in LLM Fine-tuning
Sirui Chen, Yunzhe Qi, Mengting Ai +4
Supervised fine-tuning (SFT) relies critically on selecting training data that most benefits a model's downstream performance. Gradient-based data selection methods such as TracIn…
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
Feiyang Kang, Yifan Sun, Bingbing Wen +4
Domain reweighting is an emerging research area aimed at adjusting the relative weights of different data sources to improve the effectiveness and efficiency of LLM pre-training. W…