papers

Publications (88)

cs.LG2019

Local Saddle Point Optimization: A Curvature Exploitation Approach

Leonard Adolphs, Hadi Daneshmand, Aurelien Lucchi +1

Gradient-based optimization methods are the most popular choice for finding local optima for classical minimization and saddle point problems. Here, we highlight a systemic issue o…

cs.LG2021

Scalable Graph Networks for Particle Simulations

Karolis Martinkus, Aurelien Lucchi, Nathanaël Perraudin

Learning system dynamics directly from observations is a promising direction in machine learning due to its potential to significantly enhance our ability to understand physical sy…

cs.LG2026

Adaptive Methods Are Preferable in High Privacy Settings: An SDE Perspective

Enea Monzio Compagnoni, Alessandro Stanghellini, Rustem Islamov +2

Differential Privacy (DP) is becoming central to large-scale training as privacy regulations tighten. We revisit how DP noise interacts with adaptivity in optimization through the…

cs.LG2025

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size

Rustem Islamov, Niccolo Ajroldi, Antonio Orvieto +1

Modern optimization algorithms that incorporate momentum and adaptive step-size offer improved performance in numerous challenging deep learning tasks. However, their effectiveness…

stat.ML2024

A Theoretical Analysis of the Learning Dynamics under Class Imbalance

Emanuele Francazi, Marco Baity-Jesi, Aurelien Lucchi

Data imbalance is a common problem in machine learning that can have a critical effect on the performance of a model. Various solutions exist but their impact on the convergence of…

cs.LG2020

The Role of Memory in Stochastic Optimization

Antonio Orvieto, Jonas Kohler, Aurelien Lucchi

The choice of how to retain information about past gradients dramatically affects the convergence properties of state-of-the-art stochastic optimization methods, such as Heavy-ball…

quant-ph2026

Gradient Scalability and Taylor Surrogation of Quantum Cost Landscapes

Sabri Meyer, Francesco Scala, Francesco Tacchino +1

Variational Quantum Algorithms are promising candidates for near-term quantum computing, yet they face scalability challenges due to barren plateaus, where gradients vanish exponen…

astro-ph.CO2019

Cosmological constraints with deep learning from KiDS-450 weak lensing maps

Janis Fluri, Tomasz Kacprzak, Aurelien Lucchi +4

Convolutional Neural Networks (CNN) have recently been demonstrated on synthetic data to improve upon the precision of cosmological inference. In particular they have the potential…

cs.CL2024

Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers

Sotiris Anagnostidis, Dario Pavllo, Luca Biggio +3

Autoregressive Transformers adopted in Large Language Models (LLMs) are hard to scale to long sequences. Despite several works trying to reduce their computational cost, most of LL…

astro-ph.CO2022

A Full CDM Analysis of KiDS-1000 Weak Lensing Maps using Deep Learning

Janis Fluri, Tomasz Kacprzak, Aurelien Lucchi +3

We present a full forward-modeled CDM analysis of the KiDS-1000 weak lensing maps using graph-convolutional neural networks (GCNN). Utilizing the , a novel m…

cs.LG2020

A domain agnostic measure for monitoring and evaluating GANs

Paulina Grnarova, Kfir Y Levy, Aurelien Lucchi +4

Generative Adversarial Networks (GANs) have shown remarkable results in modeling complex distributions, but their evaluation remains an unsettled issue. Evaluations are essential f…

astro-ph.IM2017

Radio frequency interference mitigation using deep convolutional neural networks

Joel Akeret, Chihway Chang, Aurelien Lucchi +1

We propose a novel approach for mitigating radio frequency interference (RFI) signals in radio data using the latest advances in deep learning. We employ a special type of Convolut…

cs.LG2026

On the Interaction of Batch Noise, Adaptivity, and Compression, under -Smoothness: An SDE Approach

Enea Monzio Compagnoni, Rustem Islamov, Frank Norbert Proske +3

Distributed stochastic optimization intertwines (i) stochastic gradient noise, (ii) communication compression, and (iii) adaptive/normalized updates. While each factor has been stu…

cs.CV2017

A Semi-supervised Framework for Image Captioning

Wenhu Chen, Aurelien Lucchi, Thomas Hofmann

State-of-the-art approaches for image captioning require supervised training data consisting of captions with paired image data. These methods are typically unable to use unsupervi…

cs.CV2020

Convolutional Generation of Textured 3D Meshes

Dario Pavllo, Graham Spinks, Thomas Hofmann +2

While recent generative models for 2D images achieve impressive visual results, they clearly lack the ability to perform 3D reasoning. This heavily restricts the degree of control…

cs.CL2016

Probabilistic Bag-Of-Hyperlinks Model for Entity Linking

Octavian-Eugen Ganea, Marina Ganea, Aurelien Lucchi +2

Many fundamental problems in natural language processing rely on determining what entities appear in a given text. Commonly referenced as entity linking, this step is a fundamental…

cs.LG2015

A Variance Reduced Stochastic Newton Method

Aurelien Lucchi, Brian McWilliams, Thomas Hofmann

Quasi-Newton methods are widely used in practise for convex loss minimization problems. These methods exhibit good empirical performance on a wide variety of tasks and enjoy super-…

cs.LG2026

Why Do We Need Warm-up? A Theoretical Perspective

Foivos Alimisis, Rustem Islamov, Aurelien Lucchi

Learning rate warm-up -- increasing the learning rate at the beginning of training -- has become a ubiquitous heuristic in modern deep learning, yet its theoretical foundations rem…

cs.LG2025

Cubic regularized subspace Newton for non-convex optimization

Jim Zhao, Aurelien Lucchi, Nikita Doikov

This paper addresses the optimization problem of minimizing non-convex continuous functions, which is relevant in the context of high-dimensional machine learning applications char…

cs.LG2025

Unbiased and Sign Compression in Distributed Learning: Comparing Noise Resilience via SDEs

Enea Monzio Compagnoni, Rustem Islamov, Frank Norbert Proske +1

Distributed methods are essential for handling machine learning pipelines comprising large-scale models and datasets. However, their benefits often come at the cost of increased co…

cs.CL2017

Leveraging Large Amounts of Weakly Supervised Data for Multi-Language Sentiment Classification

Jan Deriu, Aurelien Lucchi, Valeria De Luca +5

This paper presents a novel approach for multi-lingual sentiment classification in short texts. This is a challenging task as the amount of training data in languages other than En…

cs.CV2022

Mastering Spatial Graph Prediction of Road Networks

Sotiris Anagnostidis, Aurelien Lucchi, Thomas Hofmann

Accurately predicting road networks from satellite images requires a global understanding of the network topology. We propose to capture such high-level information by introducing…

cs.LG2021

Neural Symbolic Regression that Scales

Luca Biggio, Tommaso Bendinelli, Alexander Neitz +2

Symbolic equations are at the core of scientific discovery. The task of discovering the underlying equation from a set of input-output pairs is called symbolic regression. Traditio…

cs.LG2016

DynaNewton - Accelerating Newton's Method for Machine Learning

Hadi Daneshmand, Aurelien Lucchi, Thomas Hofmann

Newton's method is a fundamental technique in optimization with quadratic convergence within a neighborhood around the optimum. However reaching this neighborhood is often slow and…

cs.LG2017

An Online Learning Approach to Generative Adversarial Networks

Paulina Grnarova, Kfir Y. Levy, Aurelien Lucchi +2

We consider the problem of training generative models with a Generative Adversarial Network (GAN). Although GANs can accurately model complex distributions, they are known to be di…

cs.NE2022

A Globally Convergent Evolutionary Strategy for Stochastic Constrained Optimization with Applications to Reinforcement Learning

Youssef Diouane, Aurelien Lucchi, Vihang Patil

Evolutionary strategies have recently been shown to achieve competing levels of performance for complex optimization problems in reinforcement learning. In such problems, one often…

astro-ph.CO2021

Cosmological Parameter Estimation and Inference using Deep Summaries

Janis Fluri, Aurelien Lucchi, Tomasz Kacprzak +2

The ability to obtain reliable point estimates of model parameters is of crucial importance in many fields of physics. This is often a difficult task given that the observed data c…

math.OC2021

Momentum Improves Optimization on Riemannian Manifolds

Foivos Alimisis, Antonio Orvieto, Gary Bécigneul +1

We develop a new Riemannian descent algorithm that relies on momentum to improve over existing first-order methods for geodesically convex optimization. In contrast, accelerated co…

physics.comp-ph2019

Cosmological N-body simulations: a challenge for scalable generative models

Nathanaël Perraudin, Ankit Srivastava, Aurelien Lucchi +3

Deep generative models, such as Generative Adversarial Networks (GANs) or Variational Autoencoders (VAs) have been demonstrated to produce images of high visual quality. However, t…

cs.LG2024

Initial Guessing Bias: How Untrained Networks Favor Some Classes

Emanuele Francazi, Aurelien Lucchi, Marco Baity-Jesi

Understanding and controlling biasing effects in neural networks is crucial for ensuring accurate and fair model performance. In the context of classification problems, we provide…

cs.LG2018

A Distributed Second-Order Algorithm You Can Trust

Celestine Dünner, Aurelien Lucchi, Matilde Gargiani +3

Due to the rapid growth of data and computational resources, distributed optimization has become an active research area in recent years. While first-order methods seem to dominate…

stat.ML2025

Sharp Generalization Bounds for Foundation Models with Asymmetric Randomized Low-Rank Adapters

Anastasis Kratsios, Tin Sum Cheng, Aurelien Lucchi +1

Low-Rank Adaptation (LoRA) has emerged as a widely adopted parameter-efficient fine-tuning (PEFT) technique for foundation models. Recent work has highlighted an inherent asymmetry…

cs.CV2020

Controlling Style and Semantics in Weakly-Supervised Image Generation

Dario Pavllo, Aurelien Lucchi, Thomas Hofmann

We propose a weakly-supervised approach for conditional image generation of complex scenes where a user has fine control over objects appearing in the scene. We exploit sparse sema…

cs.LG2018

Escaping Saddles with Stochastic Gradients

Hadi Daneshmand, Jonas Kohler, Aurelien Lucchi +1

We analyze the variance of stochastic gradients along negative curvature directions in certain non-convex machine learning models and show that stochastic gradients exhibit a stron…

cs.LG2017

Stabilizing Training of Generative Adversarial Networks through Regularization

Kevin Roth, Aurelien Lucchi, Sebastian Nowozin +1

Deep generative models based on Generative Adversarial Networks (GANs) have demonstrated impressive sample quality but in order to work they require a careful choice of architectur…

astro-ph.CO2017

Cosmological model discrimination with Deep Learning

Jorit Schmelzle, Aurelien Lucchi, Tomasz Kacprzak +4

We demonstrate the potential of Deep Learning methods for measurements of cosmological parameters from density fields, focusing on the extraction of non-Gaussian information. We co…

cs.LG2025

Theoretical characterisation of the Gauss-Newton conditioning in Neural Networks

Jim Zhao, Sidak Pal Singh, Aurelien Lucchi

The Gauss-Newton (GN) matrix plays an important role in machine learning, most evident in its use as a preconditioning matrix for a wide family of popular adaptive methods to speed…

math.OC2023

A Sub-sampled Tensor Method for Non-convex Optimization

Aurelien Lucchi, Jonas Kohler

We present a stochastic optimization method that uses a fourth-order regularized model to find local minima of smooth and potentially non-convex objective functions with a finite-s…

cs.LG2026

Where You Place the Norm Matters: From Prejudiced to Neutral Initializations

Emanuele Francazi, Francesco Pinto, Aurelien Lucchi +1

Normalization layers were introduced to stabilize and accelerate training, yet their influence is critical already at initialization, where they shape signal propagation and output…

math.OC2021

On the Second-order Convergence Properties of Random Search Methods

Aurelien Lucchi, Antonio Orvieto, Adamos Solomou

We study the theoretical convergence properties of random-search methods when optimizing non-convex objective functions without having access to derivatives. We prove that standard…

cs.LG2026

Byzantine-Robust and Differentially Private Federated Optimization under Weaker Assumptions

Rustem Islamov, Grigory Malinovsky, Alexander Gaponov +3

Federated Learning (FL) enables heterogeneous clients to collaboratively train a shared model without centralizing their raw data, offering an inherent level of privacy. However, g…

cs.LG2026

What's in a Smoothness Constant? Tighter Rates for Local SGD with Bounded Second-order Heterogeneity

Kumar Kshitij Patel, Rustem Islamov, Sebastian U Stich +3

The paper establishes tighter convergence rates for Local SGD (Federated Averaging) on general convex problems under a bounded second‑order heterogeneity assumption, and provides n…

#local sgd#federated learning#convex optimization#heterogeneous data
cs.LG2016

Variance Reduced Stochastic Gradient Descent with Neighbors

Thomas Hofmann, Aurelien Lucchi, Simon Lacoste-Julien +1

Stochastic Gradient Descent (SGD) is a workhorse in machine learning, yet its slow convergence can be a computational bottleneck. Variance reduction techniques such as SAG, SVRG an…

cs.LG2024

A Comprehensive Analysis on the Learning Curve in Kernel Ridge Regression

Tin Sum Cheng, Aurelien Lucchi, Anastasis Kratsios +1

This paper conducts a comprehensive study of the learning curves of kernel ridge regression (KRR) under minimal assumptions. Our contributions are three-fold: 1) we analyze the rol…

cs.LG2026

When Bias Meets Trainability: Connecting Theories of Initialization

Alberto Bassi, Marco Baity-Jesi, Aurelien Lucchi +2

The statistical properties of deep neural networks (DNNs) at initialization play an important role to comprehend their trainability and the intrinsic architectural biases they poss…

cs.LG2023

A Theoretical Analysis of the Test Error of Finite-Rank Kernel Ridge Regression

Tin Sum Cheng, Aurelien Lucchi, Ivan Dokmanić +2

Existing statistical learning guarantees for general kernel regressors often yield loose bounds when used with finite-rank kernels. Yet, finite-rank kernels naturally appear in sev…

cs.LG2023

An SDE for Modeling SAM: Theory and Insights

Enea Monzio Compagnoni, Luca Biggio, Antonio Orvieto +3

We study the SAM (Sharpness-Aware Minimization) optimizer which has recently attracted a lot of interest due to its increased performance over more classical variants of stochastic…

cs.LG2016

Starting Small -- Learning with Adaptive Sample Sizes

Hadi Daneshmand, Aurelien Lucchi, Thomas Hofmann

For many machine learning problems, data is abundant and it may be prohibitive to make multiple passes through the full training set. In this context, we investigate strategies for…

stat.ML2022

Phenomenology of Double Descent in Finite-Width Neural Networks

Sidak Pal Singh, Aurelien Lucchi, Thomas Hofmann +1

`Double descent' delineates the generalization behaviour of models depending on the regime they belong to: under- or over-parameterized. The current theoretical understanding behin…

cs.LG2018

Semantic Interpolation in Implicit Models

Yannic Kilcher, Aurelien Lucchi, Thomas Hofmann

In implicit models, one often interpolates between sampled points in latent space. As we show in this paper, care needs to be taken to match-up the distributional assumptions on co…

cs.LG2022

Signal Propagation in Transformers: Theoretical Perspectives and the Role of Rank Collapse

Lorenzo Noci, Sotiris Anagnostidis, Luca Biggio +3

Transformers have achieved remarkable success in several domains, ranging from natural language processing to computer vision. Nevertheless, it has been recently shown that stackin…

math.PR2023

Mean first exit times of Ornstein-Uhlenbeck processes in high-dimensional spaces

Hans Kersting, Antonio Orvieto, Frank Proske +1

The -dimensional Ornstein--Uhlenbeck process (OUP) describes the trajectory of a particle in a -dimensional, spherically symmetric, quadratic potential. The OUP is composed o…

cs.LG2024

Characterizing Overfitting in Kernel Ridgeless Regression Through the Eigenspectrum

Tin Sum Cheng, Aurelien Lucchi, Anastasis Kratsios +1

We derive new bounds for the condition number of kernel matrices, which we then use to enhance existing non-asymptotic test error bounds for kernel ridgeless regression (KRR) in th…

cs.LG2020

Randomized Block-Diagonal Preconditioning for Parallel Learning

Celestine Mendler-Dünner, Aurelien Lucchi

We study preconditioned gradient-based optimization methods where the preconditioning matrix has block-diagonal form. Such a structural constraint comes with the advantage that the…

stat.ML2020

Batch Normalization Provably Avoids Rank Collapse for Randomly Initialised Deep Networks

Hadi Daneshmand, Jonas Kohler, Francis Bach +2

Randomly initialized neural networks are known to become harder to train with increasing depth, unless architectural enhancements like residual connections and batch normalization…

cs.CL2026

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

Shyam Sundhar Ramesh, Xiaotong Ji, Matthieu Zimmer +5

RL-based post-training with GRPO is widely used to improve large language models on individual reasoning tasks. However, real-world deployment requires reliable performance across…

cs.LG2026

Beyond a Single Explanation of the Adam--SGD Gap

Chenxiang Zhang, Rustem Islamov, Enea Monzio Compagnoni +3

Prior work has identified several factors that can contribute to the performance gap between Adam and SGD, spanning data aspects, architecture design, and optimization properties.…

cs.CV2017

Learning Aerial Image Segmentation from Online Maps

Pascal Kaiser, Jan Dirk Wegner, Aurelien Lucchi +3

This study deals with semantic segmentation of high-resolution (aerial) images where a semantic class label is assigned to each pixel via supervised classification as a basis for a…

cs.CV2021

Learning Generative Models of Textured 3D Meshes from Real-World Images

Dario Pavllo, Jonas Kohler, Thomas Hofmann +1

Recent advances in differentiable rendering have sparked an interest in learning generative models of textured 3D meshes from image collections. These models natively disentangle p…

astro-ph.CO2021

Emulation of cosmological mass maps with conditional generative adversarial networks

Nathanaël Perraudin, Sandro Marcon, Aurelien Lucchi +1

Weak gravitational lensing mass maps play a crucial role in understanding the evolution of structures in the universe and our ability to constrain cosmological models. The predicti…

cs.LG2026

Double Momentum and Error Feedback for Clipping with Fast Rates and Differential Privacy

Rustem Islamov, Samuel Horvath, Aurelien Lucchi +2

Strong Differential Privacy (DP) and Optimization guarantees are two desirable properties for a method in Federated Learning (FL). However, existing algorithms do not achieve both…

cs.LG2024

Loss Landscape Characterization of Neural Networks without Over-Parametrization

Rustem Islamov, Niccolò Ajroldi, Antonio Orvieto +1

Optimization methods play a crucial role in modern machine learning, powering the remarkable empirical achievements of deep learning models. These successes are even more remarkabl…

stat.ML2023

Anticorrelated Noise Injection for Improved Generalization

Antonio Orvieto, Hans Kersting, Frank Proske +2

Injecting artificial noise into gradient descent (GD) is commonly employed to improve the performance of machine learning models. Usually, uncorrelated noise is used in such pertur…

astro-ph.CO2018

Cosmological constraints from noisy convergence maps through deep learning

Janis Fluri, Tomasz Kacprzak, Aurelien Lucchi +3

Deep learning is a powerful analysis technique that has recently been proposed as a method to constrain cosmological parameters from weak lensing mass maps. Due to its ability to l…

cs.LG2017

Sub-sampled Cubic Regularization for Non-convex Optimization

Jonas Moritz Kohler, Aurelien Lucchi

We consider the minimization of non-convex functions that typically arise in machine learning. Specifically, we focus our attention on a variant of trust region methods known as cu…

cs.LG2025

Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise

Enea Monzio Compagnoni, Tianlin Liu, Rustem Islamov +3

Despite the vast empirical evidence supporting the efficacy of adaptive optimization methods in deep learning, their theoretical understanding is far from complete. This work intro…

cs.LG2021

Faster Single-loop Algorithms for Minimax Optimization without Strong Concavity

Junchi Yang, Antonio Orvieto, Aurelien Lucchi +1

Gradient descent ascent (GDA), the simplest single-loop algorithm for nonconvex minimax optimization, is widely used in practical applications such as generative adversarial networ…

cs.LG2021

Generative Minimization Networks: Training GANs Without Competition

Paulina Grnarova, Yannic Kilcher, Kfir Y. Levy +2

Many applications in machine learning can be framed as minimization problems and solved efficiently using gradient-based techniques. However, recent applications of generative mode…

math.OC2022

On the Theoretical Properties of Noise Correlation in Stochastic Optimization

Aurelien Lucchi, Frank Proske, Antonio Orvieto +2

Studying the properties of stochastic noise to optimize complex non-convex functions has been an active area of research in the field of machine learning. Prior work has shown that…

stat.ML2018

Exponential convergence rates for Batch Normalization: The power of length-direction decoupling in non-convex optimization

Jonas Kohler, Hadi Daneshmand, Aurelien Lucchi +3

Normalization techniques such as Batch Normalization have been applied successfully for training deep neural networks. Yet, despite its apparent empirical benefits, the reasons beh…

stat.ML2018

Adversarially Robust Training through Structured Gradient Regularization

Kevin Roth, Aurelien Lucchi, Sebastian Nowozin +1

We propose a novel data-dependent structured gradient regularizer to increase the robustness of neural networks vis-a-vis adversarial perturbations. Our regularizer can be derived…

math.OC2019

Shadowing Properties of Optimization Algorithms

Antonio Orvieto, Aurelien Lucchi

Ordinary differential equation (ODE) models of gradient-based optimization methods can provide insights into the dynamics of learning and inspire the design of new algorithms. Unfo…

cs.LG2026

On the Role of Batch Size in Stochastic Conditional Gradient Methods

Rustem Islamov, Roman Machacek, Aurelien Lucchi +3

We study the role of batch size in stochastic conditional gradient methods under a -Kurdyka-Łojasiewicz (-KL) condition. Focusing on momentum-based stochastic conditional…

cs.LG2024

SDEs for Minimax Optimization

Enea Monzio Compagnoni, Antonio Orvieto, Hans Kersting +2

Minimax optimization problems have attracted a lot of attention over the past few years, with applications ranging from economics to machine learning. While advanced optimization m…

cs.LG2026

Optimizer choice matters for the emergence of Neural Collapse

Jim Zhao, Tin Sum Cheng, Wojciech Masarczyk +1

Neural Collapse (NC) refers to the emergence of highly symmetric geometric structures in the representations of deep neural networks during the terminal phase of training. Despite…

math.OC2020

Continuous-time Models for Stochastic Optimization Algorithms

Antonio Orvieto, Aurelien Lucchi

We propose new continuous-time formulations for first-order stochastic optimization algorithms such as mini-batch gradient descent and variance-reduced methods. We exploit these co…

cs.LG2024

Regret-Optimal Federated Transfer Learning for Kernel Regression with Applications in American Option Pricing

Xuwei Yang, Anastasis Kratsios, Florian Krach +2

We propose an optimal iterative scheme for federated transfer learning, where a central planner has access to datasets for the same learning model $f_…

cs.LG2025

Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization

Wojciech Masarczyk, Mateusz Ostaszewski, Tin Sum Cheng +3

The softmax function is a fundamental building block of deep neural networks, commonly used to define output distributions in classification tasks or attention weights in transform…

astro-ph.IM2018

Fast Point Spread Function Modeling with Deep Learning

Jörg Herbel, Tomasz Kacprzak, Adam Amara +2

Modeling the Point Spread Function (PSF) of wide-field surveys is vital for many astrophysical applications and cosmological probes including weak gravitational lensing. The PSF sm…

cs.LG2025

Optimization Guarantees for Square-Root Natural-Gradient Variational Inference

Navish Kumar, Thomas Möllenhoff, Mohammad Emtiyaz Khan +1

Variational inference with natural-gradient descent often shows fast convergence in practice, but its theoretical convergence guarantees have been challenging to establish. This is…

cs.LG2018

Flexible Prior Distributions for Deep Generative Models

Yannic Kilcher, Aurelien Lucchi, Thomas Hofmann

We consider the problem of training generative models with deep neural networks as generators, i.e. to map latent codes to data points. Whereas the dominant paradigm combines simpl…

cs.LG2021

Vanishing Curvature and the Power of Adaptive Methods in Randomly Initialized Deep Networks

Antonio Orvieto, Jonas Kohler, Dario Pavllo +2

This paper revisits the so-called vanishing gradient phenomenon, which commonly occurs in deep randomly initialized neural networks. Leveraging an in-depth analysis of neural chain…

math.OC2021

Direct-Search for a Class of Stochastic Min-Max Problems

Sotiris Anagnostidis, Aurelien Lucchi, Youssef Diouane

Recent applications in machine learning have renewed the interest of the community in min-max optimization problems. While gradient-based optimization methods are widely used to so…

math.OC2020

A Continuous-time Perspective for Modeling Acceleration in Riemannian Optimization

Foivos Alimisis, Antonio Orvieto, Gary Bécigneul +1

We propose a novel second-order ODE as the continuous-time limit of a Riemannian accelerated gradient-based method on a manifold with curvature bounded from below. This ODE can be…

astro-ph.CO2018

Fast cosmic web simulations with generative adversarial networks

Andres C. Rodriguez, Tomasz Kacprzak, Aurelien Lucchi +5

Dark matter in the universe evolves through gravity to form a complex network of halos, filaments, sheets and voids, that is known as the cosmic web. Computational models of the un…

quant-ph2026

Noise-Induced Equalization in quantum learning models

Francesco Scala, Giacomo Guarnieri, Aurelien Lucchi

Quantum noise is known to strongly affect quantum computation, thus potentially limiting the performance of currently available quantum processing units. Even learning models based…

math.OC2020

An Accelerated DFO Algorithm for Finite-sum Convex Functions

Yuwen Chen, Antonio Orvieto, Aurelien Lucchi

Derivative-free optimization (DFO) has recently gained a lot of momentum in machine learning, spawning interest in the community to design faster methods for problems where gradien…

cs.LG2020

Adaptive norms for deep learning with regularized Newton methods

Jonas Kohler, Leonard Adolphs, Aurelien Lucchi

We investigate the use of regularized Newton methods with adaptive norms for optimizing neural networks. This approach can be seen as a second-order counterpart of adaptive gradien…