11 papers
ACPO: Asymmetric Credit Policy Optimization via Mode-Local Entropy Surrogate
Zijun Xie, Yuyang You, Yongzhi Li +8
Outcome-supervised reinforcement learning scales to verifiable reasoning tasks, but trajectory-level rewards assign the same outcome signal to all sampled tokens, overlooking their…
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning
Xikai Zhang, Yongzhi Li, Likang Xiao +6
Reinforcement learning has become a cornerstone for aligning and unlocking the reasoning capabilities of large-scale models. At its core, the training loop of GRPO and its variants…
Recommendation as Generation: Unifying Personalized Video Generation and Recommendation at Industrial Scale
Yanhua Cheng, Bo Wang, Haotian Zhang +17
Traditional short-video recommendation systems match user interest to a fixed pool of pre-produced videos, which limits their ability to capture fine-grained and dynamic preference…
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
MiniMax, :, Aili Chen +219
We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
Chenghao Li, Fusheng Hao, Xikai Zhang +5
Multimodal large language models via reinforcement learning (RL) have demonstrated remarkable capabilities in complex visual reasoning tasks, yet they remain limited in long-horizo…
Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing
Yanbo Wang, Yuxuan Wang, Chen Chen +6
With the wide adoption of Multimodal Models (MMs) in real-world scenarios, it is significant to efficiently train emerging MMs that exhibit increasingly complex module architecture…