Showing cs.AIShow all
3 papers · 1 filter
cs.AI2025
MAPO: Mixed Advantage Policy Optimization
Wenke Huang, Quan Zhang, Yiyang Fang +11
Recent advances in reinforcement learning for foundation models, such as Group Relative Policy Optimization (GRPO), have significantly improved the performance of foundation models…
cs.AI2025
CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process
Jinhe Bi, Danqi Yan, Yifan Wang +8
Recent Large Reasoning Models significantly improve the reasoning ability of Large Language Models by learning to reason, exhibiting the promising performance in solving complex ta…
cs.AI2025
Privacy-Enhancing Paradigms within Federated Multi-Agent Systems
Zitong Shi, Guancheng Wan, Wenke Huang +4
LLM-based Multi-Agent Systems (MAS) have proven highly effective in solving complex problems by integrating multiple agents, each performing different roles. However, in sensitive…