4 papers
SMART Fine-tuning Factor Augmented Neural Lasso
Jinhang Chai, Jianqing Fan, Cheng Gao +1
Fine-tuning is a widely used strategy for adapting pre-trained models to new tasks, yet its methodology and theoretical properties in high-dimensional nonparametric settings with v…
Robust Transfer Learning with Unreliable Source Data
Jianqing Fan, Cheng Gao, Jason M. Klusowski
This paper addresses challenges in robust transfer learning stemming from ambiguity in Bayes classifiers and weak transferable signals between the target and source distribution. W…
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
Zihao Li, Yuan Cao, Cheng Gao +5
Transformers have achieved great success in recent years. Interestingly, transformers have shown particularly strong in-context learning capability -- even without fine-tuning, the…
Global Convergence in Training Large-Scale Transformers
Cheng Gao, Yuan Cao, Zihao Li +5
Despite the widespread success of Transformers across various domains, their optimization guarantees in large-scale model settings are not well-understood. This paper rigorously an…