3 citations · 5 across the 12 of their papers we have counts for
Showing stat.MLShow all
2 papers · 1 filter
stat.ML2025
Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes
Ruiqi Zhang, Jingfeng Wu, Licong Lin +1
We study (GD) for logistic regression on linearly separable data with stepsizes that adapt to the current risk, scaled by a constant hyperparameter .…
stat.ML2024★ 1 cited
In-Context Learning of a Linear Transformer Block: Benefits of the MLP Component and One-Step GD Initialization
Ruiqi Zhang, Jingfeng Wu, Peter L. Bartlett
We study the \emph{in-context learning} (ICL) ability of a \emph{Linear Transformer Block} (LTB) that combines a linear attention component and a linear multi-layer perceptron (MLP…