9 citations · 44 across the 9 of their papers we have counts for
7 papers · 1 filter
The Law of Parsimony in Gradient Descent for Learning Deep Linear Networks
Can Yaras, Peng Wang, Wei Hu +3
Over the past few years, an extensively studied phenomenon in training deep networks is the implicit bias of gradient descent towards parsimonious solutions. In this work, we inves…
Are All Losses Created Equal: A Neural Collapse Perspective
Jinxin Zhou, Chong You, Xiao Li +4
While cross entropy (CE) is the most commonly used loss to train deep neural networks for classification tasks, many alternative losses have been developed to obtain better empiric…
On the Optimization Landscape of Neural Collapse under MSE Loss: Global Optimality with Unconstrained Features
Jinxin Zhou, Xiao Li, Tianyu Ding +3
When training deep neural networks for classification tasks, an intriguing empirical phenomenon has been widely observed in the last-layer classifiers and features, where (i) the c…
A Geometric Analysis of Neural Collapse with Unconstrained Features
Zhihui Zhu, Tianyu Ding, Jinxin Zhou +4
We provide the first global optimization landscape analysis of -- an intriguing empirical phenomenon that arises in the last-layer classifiers and features of ne…
Robust Recovery via Implicit Bias of Discrepant Learning Rates for Double Over-parameterization
Chong You, Zhihui Zhu, Qing Qu +1
Recent advances have shown that implicit bias of gradient descent on over-parameterized models enables the recovery of low-rank matrices from linear measurements, even with no prio…
Finding the Sparsest Vectors in a Subspace: Theory, Algorithms, and Applications
Qing Qu, Zhihui Zhu, Xiao Li +3
The problem of finding the sparsest vector (direction) in a low dimensional subspace can be considered as a homogeneous variant of the sparse recovery problem, which finds applicat…