3 citations · 12 across the 10 of their papers we have counts for
10 papers
Learning Scalable Model Soup on a Single GPU: An Efficient Subspace Training Strategy
Tao Li, Weisen Jiang, Fanghui Liu +2
Pre-training followed by fine-tuning is widely adopted among practitioners. The performance can be improved by "model soups"~\cite{wortsman2022model} via exploring various hyperpar…
High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit Regularization
Yihang Chen, Fanghui Liu, Taiji Suzuki +1
This paper studies kernel ridge regression in high dimensions under covariate shifts and analyzes the role of importance re-weighting. We first derive the asymptotic expansion of h…
Robust NAS under adversarial training: benchmark, theory, and beyond
Yongtao Wu, Fanghui Liu, Carl-Johann Simon-Gabriel +2
Recent developments in neural architecture search (NAS) emphasize the significance of considering robust architectures against malicious data. However, there is a notable absence o…
Generalization of Scaled Deep ResNets in the Mean-Field Regime
Yihang Chen, Fanghui Liu, Yiping Lu +2
Despite the widespread empirical success of ResNet, the generalization properties of deep ResNet are rarely explored beyond the lazy training regime. In this work, we investigate \…
Efficient local linearity regularization to overcome catastrophic overfitting
Elias Abad Rocamora, Fanghui Liu, Grigorios G. Chrysos +2
Catastrophic overfitting (CO) in single-step adversarial training (AT) results in abrupt drops in the adversarial test accuracy (even down to 0%). For models trained with multi-ste…
On the Convergence of Encoder-only Shallow Transformers
Yongtao Wu, Fanghui Liu, Grigorios G Chrysos +1
In this paper, we aim to build the global convergence theory of encoder-only shallow Transformers under a realistic setting from the perspective of architectures, initialization, a…