3 papers
math.OC2026
Smooth or Separable? Sparse Block Acceleration for Entropy-Regularized Linear Programs
Irina Podlipnova, Maxim Mashtaler, Artem Agafonov +2
We study entropy-regularized linear programs with sparse affine constraints and compare two exact dual representations induced by whether a redundant normalization constraint is re…
math.OC2026
Application of Optimal Inexact Second-Order Acceleration to Distributed Stochastic Optimization under Statistical Similarity
Yury A. Sokolov, Maxim K. Mashtaler, Alexander V. Gasnikov +3
We consider distributed stochastic convex optimization with a fixed budget of independent samples split among workers. Sample average approximation reduces the problem to a…
cs.LG2025
AdLoCo: adaptive batching significantly improves communications efficiency and convergence for Large Language Models
Nikolay Kutuzov, Makar Baderko, Stepan Kulibaba +4
Scaling distributed training of Large Language Models (LLMs) requires not only algorithmic advances but also efficient utilization of heterogeneous hardware resources. While existi…