4 papers
DADA: Dual Averaging with Distance Adaptation
Mohammad Moshtaghifar, Anton Rodomanov, Daniil Vankov +1
We present a novel universal gradient method for solving convex optimization problems. Our algorithm, Dual Averaging with Distance Adaptation (DADA), is based on the classical sche…
XShare: Collaborative in-Batch Expert Sharing for Faster MoE Inference
Daniil Vankov, Nikita Ivkin, Kyle Ulrich +3
Mixture-of-Experts (MoE) architectures are increasingly used to efficiently scale large language models. However, in production inference, request batching and speculative decoding…
Generalized Smooth Stochastic Variational Inequalities: Almost Sure Convergence and Convergence Rates
Daniil Vankov, Angelia Nedich, Lalitha Sankar
This paper focuses on solving a stochastic variational inequality (SVI) problem under relaxed smoothness assumption for a class of structured non-monotone operators. The SVI proble…
Optimizing -Smooth Functions by Gradient Methods
Daniil Vankov, Anton Rodomanov, Angelia Nedich +2
We study gradient methods for optimizing -smooth functions, a class that generalizes Lipschitz-smooth functions and has gained attention for its relevance in machine le…