most citedA Unified Analysis for Finite Weight Averaging

1 citations · 1 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CL2025

Joint Selection for Large-Scale Pre-Training Data via Policy Gradient-based Mask Learning

Ziqing Fan, Yuqiao Xian, Yan Sun +1

A fine-grained data recipe is crucial for pre-training large language models, as it can significantly enhance training efficiency and model performance. One important ingredient in…

cs.MA2025

Effective Policy Learning for Multi-Agent Online Coordination Beyond Submodular Objectives

Qixin Zhang, Yan Sun, Can Jin +5

In this paper, we present two effective policy learning algorithms for multi-agent online coordination(MA-OC) problem. The first one, \texttt{MA-SPL}, not only can achieve the opti…

cs.LG2025

A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Zihao Lin, Samyadeep Basu, Mohammad Beigi +18

The rise of foundation models has transformed machine learning research, prompting efforts to uncover their inner workings and develop more efficient and reliable applications for…

cs.LG2025

TeZO: Empowering the Low-Rankness on the Temporal Dimension in the Zeroth-Order Optimization for Fine-tuning LLMs

Yan Sun, Tiansheng Huang, Liang Ding +2

Zeroth-order optimization (ZO) has demonstrated remarkable promise in efficient fine-tuning tasks for Large Language Models (LLMs). In particular, recent advances incorporate the l…

cs.LG20241 cited

A Unified Analysis for Finite Weight Averaging

Peng Wang, Li Shen, Zerui Tao +3

Averaging iterations of Stochastic Gradient Descent (SGD) have achieved empirical success in training deep learning models, such as Stochastic Weight Averaging (SWA), Exponential M…

cs.LG2024

Stability and Generalization for Distributed SGDA

Miaoxi Zhu, Yan Sun, Li Shen +2

Minimax optimization is gaining increasing attention in modern machine learning applications. Driven by large-scale models and massive volumes of data collected from edge devices,…