2 papers
cs.LG2025
On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning
Thomas T. Zhang, Behrad Moniri, Ansh Nagwekar +4
Layer-wise preconditioning methods are a family of memory-efficient optimization algorithms that introduce preconditioners per axis of each layer's weight tensors. These methods ha…
cs.LG2021
Rate-Distortion Analysis of Minimum Excess Risk in Bayesian Learning
Hassan Hafez-Kolahi, Behrad Moniri, Shohreh Kasaei +1
In parametric Bayesian learning, a prior is assumed on the parameter which determines the distribution of samples. In this setting, Minimum Excess Risk (MER) is defined as the…