1 citations · 1 across the 5 of their papers we have counts for
6 papers
Joint Selection for Large-Scale Pre-Training Data via Policy Gradient-based Mask Learning
Ziqing Fan, Yuqiao Xian, Yan Sun +1
A fine-grained data recipe is crucial for pre-training large language models, as it can significantly enhance training efficiency and model performance. One important ingredient in…
Effective Policy Learning for Multi-Agent Online Coordination Beyond Submodular Objectives
Qixin Zhang, Yan Sun, Can Jin +5
In this paper, we present two effective policy learning algorithms for multi-agent online coordination(MA-OC) problem. The first one, \texttt{MA-SPL}, not only can achieve the opti…
A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models
Zihao Lin, Samyadeep Basu, Mohammad Beigi +18
The rise of foundation models has transformed machine learning research, prompting efforts to uncover their inner workings and develop more efficient and reliable applications for…
TeZO: Empowering the Low-Rankness on the Temporal Dimension in the Zeroth-Order Optimization for Fine-tuning LLMs
Yan Sun, Tiansheng Huang, Liang Ding +2
Zeroth-order optimization (ZO) has demonstrated remarkable promise in efficient fine-tuning tasks for Large Language Models (LLMs). In particular, recent advances incorporate the l…
A Unified Analysis for Finite Weight Averaging
Peng Wang, Li Shen, Zerui Tao +3
Averaging iterations of Stochastic Gradient Descent (SGD) have achieved empirical success in training deep learning models, such as Stochastic Weight Averaging (SWA), Exponential M…
Stability and Generalization for Distributed SGDA
Miaoxi Zhu, Yan Sun, Li Shen +2
Minimax optimization is gaining increasing attention in modern machine learning applications. Driven by large-scale models and massive volumes of data collected from edge devices,…