Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Zeta: Dual Whitening for Matrix Optimization via Coordinate-Adaptive Preconditioning
Kaiwen Chen, Shuhai Zhang, Zimo Liu +7
Large-scale neural network training increasingly relies on matrix-aware optimizers that exploit the structure of weight parameters beyond element-wise adaptation. However, existing…
cs.LG2023
Mixture Data for Training Cannot Ensure Out-of-distribution Generalization
Songming Zhang, Yuxiao Luo, Qizhou Wang +4
Deep neural networks often face generalization problems to handle out-of-distribution (OOD) data, and there remains a notable theoretical gap between the contributing factors and t…