works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.LG2026

Gradient Flow Dynamics and Implicit Bias of Diagonal Linear Networks under Infinitesimal Initialization

Jiajie Zhao, Jianxing Wang, Junjie Yang +2

The paper analyzes how diagonal linear networks trained with infinitesimal initialization evolve under gradient flow, showing they follow a specific algorithm that converges to a m…

cs.LG2026

Towards Understanding Adam Convergence on Highly Degenerate Polynomials

Zhiwei Bai, Jiajie Zhao, Zhangchen Zhou +2

Adam is a widely used optimization algorithm in deep learning, yet the specific class of objective functions where it exhibits inherent advantages remains underexplored. Unlike pri…

cs.LG2026

Adaptive Preconditioners Trigger Loss Spikes in Adam

Zhiwei Bai, Zhangchen Zhou, Jiajie Zhao +6

Loss spikes commonly emerge during neural network training with the Adam optimizer across diverse architectures and scales, yet their underlying mechanism remains elusive. While pr…

cs.LG2025

Scalable Complexity Control Facilitates Reasoning Ability of LLMs

Liangkai Hang, Junjie Yao, Zhiwei Bai +17

The reasoning ability of large language models (LLMs) has been rapidly advancing in recent years, attracting interest in more fundamental approaches that can reliably enhance their…

cs.LG2025

Connectivity Shapes Implicit Regularization in Matrix Factorization Models for Matrix Completion

Zhiwei Bai, Jiajie Zhao, Yaoyu Zhang

Matrix factorization models have been extensively studied as a valuable test-bed for understanding the implicit biases of overparameterized models. Although both low nuclear norm a…