collaborators

6 papers

cs.AI2026

MiniMax Sparse Attention

Xunhao Lai, Weiqi Xu, Yufeng Yang +14

Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointl…

cs.AI2026

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

MiniMax, :, Aili Chen +219

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…

cs.LG2026

When Correct Demonstrations Hurt: Rethinking the Role of Exemplars in In-Context Learning

Chenghao Qiu, Chunli Peng, Yufeng Yang +2

In-context learning (ICL) is often motivated by the intuition that demonstrations help because they provide correct input-output examples. However, we reveal a counterintuitive phe…

cs.LG2026

Distributionally Robust Multi-Objective Optimization

Yufeng Yang, Fangning Zhuo, Ziyi Chen +2

Multi-objective optimization (MOO) has received growing attention in applications that require learning under multiple criteria. However, the existing MOO formulations do not expli…

math.OC2025

Adaptive Gradient Normalization and Independent Sampling for (Stochastic) Generalized-Smooth Optimization

Yufeng Yang, Erin Tripp, Yifan Sun +2

Recent studies have shown that many nonconvex machine learning problems satisfy a generalized-smooth condition that extends beyond traditional smooth nonconvex optimization. Howeve…

math.OC2025

Nested Stochastic Algorithm for Generalized Sinkhorn distance-Regularized Distributionally Robust Optimization

Yufeng Yang, Yi Zhou, Zhaosong Lu

Distributionally robust optimization (DRO) is a powerful technique to train robust models against data distribution shift. This paper aims to solve regularized nonconvex DRO proble…