3 papers
cs.LG2026
Exploiting weight-space symmetries for approximating curvature
Artem Artemev, Rui Xia, Benjamin M. Boyd +4
Many machine learning techniques rely on approximating a loss function's curvature, but this is notoriously hard to do at the scale of modern deep networks. Surprisingly, no previo…
cs.LG2025
Perturbation-efficient Zeroth-order Optimization for Hardware-friendly On-device Training
Qitao Tan, Sung-En Chang, Rui Xia +10
Zeroth-order (ZO) optimization is an emerging deep neural network (DNN) training paradigm that offers computational simplicity and memory savings. However, this seemingly promising…
cs.LG2024
Efficient Model Compression Techniques with FishLeg
Jamie McGowan, Wei Sheng Lai, Weibin Chen +7
In many domains, the most successful AI models tend to be the largest, indeed often too large to be handled by AI players with limited computational resources. To mitigate this, a…