3 papers
cs.AI2026
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
MiniMax, :, Aili Chen +219
We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…
math.OC2025
Convergence Analysis of Stochastic Accelerated Gradient Methods for Generalized Smooth Optimizations
Chenhao Yu, Yusu Hong, Junhong Lin
We investigate the Randomized Stochastic Accelerated Gradient (RSAG) method, utilizing either constant or adaptive step sizes, for stochastic optimization problems with generalized…
math.OC2025
On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions
Yusu Hong, Junhong Lin
The Adaptive Momentum Estimation (Adam) algorithm is highly effective in training various deep learning tasks. Despite this, there's limited theoretical understanding for Adam, esp…