6 papers
MiniMax Sparse Attention
Xunhao Lai, Weiqi Xu, Yufeng Yang +14
Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointl…
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
MiniMax, :, Aili Chen +219
We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…
When Correct Demonstrations Hurt: Rethinking the Role of Exemplars in In-Context Learning
Chenghao Qiu, Chunli Peng, Yufeng Yang +2
In-context learning (ICL) is often motivated by the intuition that demonstrations help because they provide correct input-output examples. However, we reveal a counterintuitive phe…
Distributionally Robust Multi-Objective Optimization
Yufeng Yang, Fangning Zhuo, Ziyi Chen +2
Multi-objective optimization (MOO) has received growing attention in applications that require learning under multiple criteria. However, the existing MOO formulations do not expli…
Adaptive Gradient Normalization and Independent Sampling for (Stochastic) Generalized-Smooth Optimization
Yufeng Yang, Erin Tripp, Yifan Sun +2
Recent studies have shown that many nonconvex machine learning problems satisfy a generalized-smooth condition that extends beyond traditional smooth nonconvex optimization. Howeve…
Nested Stochastic Algorithm for Generalized Sinkhorn distance-Regularized Distributionally Robust Optimization
Yufeng Yang, Yi Zhou, Zhaosong Lu
Distributionally robust optimization (DRO) is a powerful technique to train robust models against data distribution shift. This paper aims to solve regularized nonconvex DRO proble…