activity
20242026
collaborators

6 papers

cs.AI2026

MiniMax Sparse Attention

Xunhao Lai, Weiqi Xu, Yufeng Yang +14

Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointl…

cs.LG2026

When Correct Demonstrations Hurt: Rethinking the Role of Exemplars in In-Context Learning

Chenghao Qiu, Chunli Peng, Yufeng Yang +2

In-context learning (ICL) is often motivated by the intuition that demonstrations help because they provide correct input-output examples. However, we reveal a counterintuitive phe…

cs.LG2026

Distributionally Robust Multi-Objective Optimization

Yufeng Yang, Fangning Zhuo, Ziyi Chen +2

Multi-objective optimization (MOO) has received growing attention in applications that require learning under multiple criteria. However, the existing MOO formulations do not expli…

cs.AI2026

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

Aili Chen, Aonian Li, Baichuan Zhou +215

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…

math.OC2025

Nested Stochastic Algorithm for Generalized Sinkhorn distance-Regularized Distributionally Robust Optimization

Yufeng Yang, Yi Zhou, Zhaosong Lu

Distributionally robust optimization (DRO) is a powerful technique to train robust models against data distribution shift. This paper aims to solve regularized nonconvex DRO proble…

math.OC2024

Adaptive Gradient Normalization and Independent Sampling for (Stochastic) Generalized-Smooth Optimization

Yufeng Yang, Erin Tripp, Yifan Sun +2

Recent studies have shown that many nonconvex machine learning problems satisfy a generalized-smooth condition that extends beyond traditional smooth nonconvex optimization. Howeve…