6 papers
MiniMax Sparse Attention
Xunhao Lai, Weiqi Xu, Yufeng Yang +14
Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointl…
When Correct Demonstrations Hurt: Rethinking the Role of Exemplars in In-Context Learning
Chenghao Qiu, Chunli Peng, Yufeng Yang +2
In-context learning (ICL) is often motivated by the intuition that demonstrations help because they provide correct input-output examples. However, we reveal a counterintuitive phe…
Distributionally Robust Multi-Objective Optimization
Yufeng Yang, Fangning Zhuo, Ziyi Chen +2
Multi-objective optimization (MOO) has received growing attention in applications that require learning under multiple criteria. However, the existing MOO formulations do not expli…
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
Aili Chen, Aonian Li, Baichuan Zhou +215
We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…
Nested Stochastic Algorithm for Generalized Sinkhorn distance-Regularized Distributionally Robust Optimization
Yufeng Yang, Yi Zhou, Zhaosong Lu
Distributionally robust optimization (DRO) is a powerful technique to train robust models against data distribution shift. This paper aims to solve regularized nonconvex DRO proble…
Adaptive Gradient Normalization and Independent Sampling for (Stochastic) Generalized-Smooth Optimization
Yufeng Yang, Erin Tripp, Yifan Sun +2
Recent studies have shown that many nonconvex machine learning problems satisfy a generalized-smooth condition that extends beyond traditional smooth nonconvex optimization. Howeve…