2 papers
cs.AI2026
StaRPO: Stability-Augmented Reinforcement Policy Optimization
Jinghan Zhang, Fengran Mo, Tharindu Cyril Weerasooriya +5
Reinforcement learning (RL) is effective in enhancing the accuracy of large language models in complex reasoning tasks. Existing RL policy optimization frameworks rely on final-ans…
cs.AI2025
Galaxy: A Cognition-Centered Framework for Proactive, Privacy-Preserving, and Self-Evolving LLM Agents
Chongyu Bao, Ruimin Dai, Yangbo Shen +4
Intelligent personal assistants (IPAs) such as Siri and Google Assistant are designed to enhance human capabilities and perform tasks on behalf of users. The emergence of LLM agent…