1 paper
Shih-Yu Sun, Vimal Thilak, Etai Littwin +2
Deep linear networks trained with gradient descent yield low rank solutions, as is typically studied in matrix factorization. In this paper, we take a step further and analyze impl…