148 citations · 685 across the 135 of their papers we have counts for
9 papers · 2 filters
Exploring Sparsity in Graph Transformers
Chuang Liu, Yibing Zhan, Xueqi Ma +5
Graph Transformers (GTs) have achieved impressive results on various graph-related tasks. However, the huge computational cost of GTs hinders their deployment and application, espe…
Concrete Subspace Learning based Interference Elimination for Multi-task Model Fusion
Anke Tang, Xianglin Luo, Li Shen +5
Merging models fine-tuned from a common, extensively pre-trained large model but specialized for different tasks has been demonstrated as a cheap and scalable strategy to construct…
Deep Model Fusion: A Survey
Weishi Li, Yong Peng, Miao Zhang +3
Deep model fusion/merging is an emerging technique that merges the parameters or predictions of multiple deep learning models into a single one. It combines the abilities of differ…
Efficient Federated Learning via Local Adaptive Amended Optimizer with Linear Speedup
Yan Sun, Li Shen, Hao Sun +2
Adaptive optimization has achieved notable success for distributed learning while extending adaptive optimizer to federated Learning (FL) suffers from severe inefficiency, includin…
Dynamic Regularized Sharpness Aware Minimization in Federated Learning: Approaching Global Consistency and Smooth Landscape
Yan Sun, Li Shen, Shixiang Chen +2
In federated learning (FL), a cluster of local clients are chaired under the coordination of the global server and cooperatively train one model with privacy protection. Due to the…
On Efficient Training of Large-Scale Deep Learning Models: A Literature Review
Li Shen, Yan Sun, Zhiyuan Yu +3
The field of deep learning has witnessed significant progress, particularly in computer vision (CV), natural language processing (NLP), and speech. The use of large-scale models tr…