actor-critic 1average-reward MDP 1distributional robustness 1q-learning 1robust reinforcement learning 1
From the 1 of 15 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
MiniMax, :, Aili Chen +219
We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…
cs.AI2026
AGPO: Asymmetric Group Policy Optimization for Verifiable Reasoning and Search Ads Relevance at JD
Yang Xu, Kun Yao, Yiming Deng +3
Reinforcement Learning with Verifiable Rewards (RLVR) has demonstrated notable success in enhancing the reasoning performance of large language models (LLMs). However, recent studi…