5 citations · 8 across the 12 of their papers we have counts for
21 papers · 1 filter
Rethinking the Personalized Relaxed Initialization in the Federated Learning: Consistency and Generalization
Li Shen, Yan Sun, Dacheng Tao
Federated learning (FL) is a distributed paradigm that coordinates massive local clients to collaboratively train a global model via stage-wise local training processes on the hete…
Multinoulli Extension: A Lossless Continuous Relaxation for Partition-Constrained Subset Selection
Qixin Zhang, Wei Huang, Yan Sun +3
Identifying the most representative subset for a close-to-submodular objective while satisfying the predefined partition constraint is a fundamental task with numerous applications…
MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs
Yan Sun, Qixin Zhang, Zhiyuan Yu +3
The rapid scaling of large language models~(LLMs) has made inference efficiency a primary bottleneck in the practical deployment. To address this, semi-structured sparsity offers a…
A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models
Zihao Lin, Samyadeep Basu, Mohammad Beigi +18
The rise of foundation models has transformed machine learning research, prompting efforts to uncover their inner workings and develop more efficient and reliable applications for…
TeZO: Empowering the Low-Rankness on the Temporal Dimension in the Zeroth-Order Optimization for Fine-tuning LLMs
Yan Sun, Tiansheng Huang, Liang Ding +2
Zeroth-order optimization (ZO) has demonstrated remarkable promise in efficient fine-tuning tasks for Large Language Models (LLMs). In particular, recent advances incorporate the l…
A Unified Analysis for Finite Weight Averaging
Peng Wang, Li Shen, Zerui Tao +3
Averaging iterations of Stochastic Gradient Descent (SGD) have achieved empirical success in training deep learning models, such as Stochastic Weight Averaging (SWA), Exponential M…