5 papers
Prism-GRPO: Faster VLA Policy Optimization via Splitting Same-outcome Groups
Zeyun Deng, Yuzhe Lu, Yawei Wang +6
GRPO is increasingly used for reinforcement learning of vision-language-action (VLA) policies because, unlike PPO, it does not require training a critic. This simplification comes…
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
Aili Chen, Aonian Li, Baichuan Zhou +215
We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…
CLEAR: Context Augmentation from Contrastive Learning of Experience via Agentic Reflection
Linbo Liu, Guande Wu, Han Ding +7
Large language model agents rely on effective model context to obtain task-relevant information for decision-making. Many existing context engineering approaches primarily rely on…
Large Language Model Agent in Financial Trading: A Survey
Han Ding, Yinheng Li, Junhao Wang +3
Trading is a highly competitive task that requires a combination of strategy, knowledge, and psychological fortitude. With the recent success of large language models(LLMs), it is…
Data Processing Techniques for Modern Multimodal Models
Yinheng Li, Han Ding, Hang Chen
Data processing plays an significant role in current multimodal model training. In this paper. we provide an comprehensive review of common data processing techniques used in moder…