5 papers
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
Sharan Sahu, Cameron J. Hogan, Martin T. Wells
In this paper, we provide a comprehensive theoretical analysis of Stochastic Gradient Descent (SGD) and its momentum variants (Polyak Heavy-Ball and Nesterov) for tracking time-var…
Modeling Dynamic Correlation Matrices with Shrinkage Priors
Daniel Andrew Coulson, David S. Matteson, Martin T. Wells
Estimating time-varying correlation matrices is challenging because existing methods may adapt slowly to structural changes, impose insufficient regularization, or produce diffuse…
Adapt or Forget: Provable Tradeoffs Between Adam and SGD in Nonstationary Optimization
Sharan Sahu, Abir Sarkar, Cameron J. Hogan +1
We provide a theoretical analysis of Adam under non-stationary stochastic objectives, separating two regimes: Euclidean tracking under adaptive strong monotonicity of the Adam-prec…
Online Distributionally Robust LLM Alignment via Regression to Relative Reward
Sharan Sahu, Martin T. Wells
Reinforcement Learning with Human Feedback (RLHF) has become crucial for aligning Large Language Models (LLMs) with human intent. However, existing offline RLHF approaches suffer f…
Minimaxity and Admissibility of Bayesian Neural Networks
Daniel Andrew Coulson, Martin T. Wells
Bayesian neural networks (BNNs) offer a natural probabilistic formulation for inference in deep learning models. Despite their popularity, their optimality has received limited att…