2 papers
cs.LG2026
Zero-Order Optimization for LLM Fine-Tuning via Learnable Direction Sampling
Valery Parfenov, Grigoriy Evseev, Andrey Veprikov +3
Fine-tuning large pretrained language models (LLMs) is a cornerstone of modern NLP, yet its growing memory demands (driven by backpropagation and large optimizer States) limit depl…
math.OC2025
Unified Theory of Adaptive Variance Reduction
Aleksandr Shestakov, Valery Parfenov, Aleksandr Beznosikov
Variance reduction is a family of powerful mechanisms for stochastic optimization that appears to be helpful in many machine learning tasks. It is based on estimating the exact gra…