activity
20142025
most citedGlobal Convergence of Online Limited Memory BFGS

131 citations · 142 across the 11 of their papers we have counts for

collaborators

11 papers

cs.LG2025

Online Learning-guided Learning Rate Adaptation via Gradient Alignment

Ruichen Jiang, Ali Kavis, Aryan Mokhtari

The performance of an optimizer on large-scale deep learning models depends critically on fine-tuning the learning rate, often requiring an extensive grid search over base learning…

math.OC2024

Improved Complexity for Smooth Nonconvex Optimization: A Two-Level Online Learning Approach with Quasi-Newton Methods

Ruichen Jiang, Aryan Mokhtari, Francisco Patitucci

We study the problem of finding an -first-order stationary point (FOSP) of a smooth function, given access only to gradient information. The best-known gradient query complexity…

math.OC2024

Krylov Cubic Regularized Newton: A Subspace Second-Order Method with Dimension-Free Convergence Rate

Ruichen Jiang, Parameswaran Raman, Shoham Sabach +3

Second-order optimization methods, such as cubic regularized Newton methods, are known for their rapid convergence rates; nevertheless, they become impractical in high-dimensional…

math.OC2023

Projection-Free Methods for Stochastic Simple Bilevel Optimization with Convex Lower-level Problem

Jincheng Cao, Ruichen Jiang, Nazanin Abolfazli +2

In this paper, we study a class of stochastic bilevel optimization problems, also known as stochastic simple bilevel optimization, where we minimize a smooth stochastic objective f…

math.OC2023

Accelerated Quasi-Newton Proximal Extragradient: Faster Rate for Smooth Convex Optimization

Ruichen Jiang, Aryan Mokhtari

In this paper, we propose an accelerated quasi-Newton proximal extragradient (A-QPNE) method for solving unconstrained smooth convex optimization problems. With access only to the…

cs.LG20231 cited

Greedy Pruning with Group Lasso Provably Generalizes for Matrix Sensing

Nived Rajaraman, Devvrit, Aryan Mokhtari +1

Pruning schemes have been widely used in practice to reduce the complexity of trained models with a massive number of parameters. In fact, several practical studies have shown that…