Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
AGPO: Asymmetric Group Policy Optimization for Verifiable Reasoning and Search Ads Relevance at JD
Yang Xu, Kun Yao, Yiming Deng +3
Reinforcement Learning with Verifiable Rewards (RLVR) has demonstrated notable success in enhancing the reasoning performance of large language models (LLMs). However, recent studi…
cs.AI2026
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
Aili Chen, Aonian Li, Baichuan Zhou +215
We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…