collaborators

7 papers

cs.LG2026

Looped Transformers with Layer Normalization Provably Learn the Power Method

Lyumin Wu, Chenyang Zhang, Yuan Cao

Transformers have achieved remarkable success across a wide range of applications, and a growing body of work suggests that part of their strength comes from their ability to learn…

cs.LG2026

Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent

Chenyang Zhang, Yuan Cao

Transformers have demonstrated remarkable in-context learning (ICL) capabilities. The strong ICL performance of transformers is commonly believed to arise from their ability to imp…

cs.LG2026

Transformers Trained via Gradient Descent Can Provably Learn a Class of Teacher Models

Chenyang Zhang, Qingyue Zhao, Quanquan Gu +1

Transformers have achieved great success across a wide range of applications, yet the theoretical foundations underlying their success remain largely unexplored. To demystify the s…

stat.ML2026

Towards Understanding Generalization in DP-GD: A Case Study in Training Two-Layer CNNs

Zhongjie Shi, Puyu Wang, Chenyang Zhang +1

Modern deep learning techniques focus on extracting intricate information from data to achieve accurate predictions. However, the training datasets may be crowdsourced and include…

stat.ML2025

Transformer Learns Optimal Variable Selection in Group-Sparse Classification

Chenyang Zhang, Xuran Meng, Yuan Cao

Transformers have demonstrated remarkable success across various applications. However, the success of transformers have not been understood in theory. In this work, we give a case…

stat.ML2025

Gradient Descent Robustly Learns the Intrinsic Dimension of Data in Training Convolutional Neural Networks

Chenyang Zhang, Peifeng Gao, Difan Zou +1

Modern neural networks are usually highly over-parameterized. Behind the wide usage of over-parameterized networks is the belief that, if the data are simple, then the trained netw…