6 papers
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
Andrey Veprikov, Arman Bolatov, Aleksandr Bogdanov +4
Optimization lies at the core of modern deep learning, yet existing methods often face a fundamental trade-off between adapting to problem geometry and leveraging curvature utiliza…
Loss-Transformation Invariance in the Damped Newton Method
Alexander Shestakov, Sushil Bohara, Samuel Horváth +2
The Newton method is a powerful optimization algorithm, valued for its rapid local convergence and elegant geometric properties. However, its theoretical guarantees are usually lim…
Simple Stepsize for Quasi-Newton Methods with Global Convergence Guarantees
Artem Agafonov, Vladislav Ryspayev, Samuel Horváth +3
Quasi-Newton methods are widely used for solving convex optimization problems due to their ease of implementation, practical efficiency, and strong local convergence guarantees. Ho…
Polyak Stepsize: Estimating Optimal Functional Values Without Parameters or Prior Knowledge
Farshed Abdukhakimov, Cuong Anh Pham, Samuel Horváth +2
The Polyak stepsize for Gradient Descent is known for its fast convergence but requires prior knowledge of the optimal functional value, which is often unavailable in practice. In…
Newton Method Revisited: Global Convergence Rates up to for Stepsize Schedules and Linesearch Procedures
SlavomÃr Hanzely, Farshed Abdukhakimov, Martin TakáÄ
This paper investigates the global convergence of stepsized Newton methods for convex functions with Hölder continuous Hessians or third derivatives. We propose several simple ste…
DAG: Projected Stochastic Approximation Iteration for DAG Structure Learning
Klea Ziu, SlavomÃr Hanzely, Loka Li +3
Learning the structure of Directed Acyclic Graphs (DAGs) presents a significant challenge due to the vast combinatorial search space of possible graphs, which scales exponentially…