5 papers
Chain-based Distillation for Effective Initialization of Variable-Sized Small Language Models
Boyu Shi, YiCheng Jiang, Chang Liu +3
Large language models (LLMs) achieve strong performance but remain costly to deploy in resource-constrained settings. Training small language models (SLMs) from scratch is computat…
Learngene Search Across Multiple Datasets for Building Variable-Sized Models
Boyu Shi, Junbo Zhou, Chang Liu +3
Deep learning methods are widely used under diverse resource constraints, resulting in models of varying sizes, such as the Vision Transformer (ViT) series. Deploying these models…
Towards Understanding Feature Learning in Parameter Transfer
Hua Yuan, Xuran Meng, Qiufeng Wang +6
Parameter transfer is a central paradigm in transfer learning, enabling knowledge reuse across tasks and domains by sharing model parameters between upstream and downstream models.…
GENE-FL: Gene-Driven Parameter-Efficient Dynamic Federated Learning
Shunxin Guo, Jiaqi Lv, Qiufeng Wang +1
Real-world \underline{F}ederated \underline{L}earning systems often encounter \underline{D}ynamic clients with \underline{A}gnostic and highly heterogeneous data distributions (DAF…
MSWAL: 3D Multi-class Segmentation of Whole Abdominal Lesions Dataset
Zhaodong Wu, Qiaochu Zhao, Ming Hu +13
With the significantly increasing incidence and prevalence of abdominal diseases, there is a need to embrace greater use of new innovations and technology for the diagnosis and tre…