2 papers
cs.LG2026
Towards On-Policy SFT: Distribution Discriminant Theory and its Applications in LLM Training
Miaosen Zhang, Yishan Liu, Shuxia Lin +8
Supervised fine-tuning (SFT) is computationally efficient but often yields inferior generalization compared to reinforcement learning (RL). This gap is primarily driven by RL's use…
cs.LG2025
STHFL: Spatio-Temporal Heterogeneous Federated Learning
Shunxin Guo, Hongsong Wang, Shuxia Lin +2
Federated learning is a new framework that protects data privacy and allows multiple devices to cooperate in training machine learning models. Previous studies have proposed multip…