From the 1 of 13 linked papers with an AI index.
13 papers
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance
Zhuowen Han, Jinwei Xiao, Zhengxi Lu +9
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models (LLMs). While Group Relative Policy Optimization (GRPO)…
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution
Zhiyuan Yao, Yuxin Chen, Zhengxi Lu +13
SkillRise introduces a reinforcement‑learning framework that lets large language model agents learn and reuse transferable skills across related tasks by curating a skill document…
Finding the Evidence: Discovering Decision-Supporting Tokens for On-Policy Reasoning Distillation
Jinwei Xiao, Zhuowen Han, Yueqing Sun +6
On-policy distillation transfers reasoning ability through dense token-level supervision, yet the nature of the transferable signal remains unclear. We discover that reasoning chai…
Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments
Yuxin Chen, Xiaodong Cai, Junfeng Fang +9
Recent advances in large language models (LLMs) have facilitated the widespread deployment of LLMs as interactive agents capable of reasoning, planning, and tool use. Despite stron…
Emotional intelligence in large language models is fragmented across perception, cognition, and interaction
Minghao Lv, Lu Chen, Enchang Zhang +6
As large language models (LLMs) are increasingly integrated into emotionally sensitive domains, the structural integrity of their emotional intelligence (EI) becomes a critical fro…
Self-Distilled Agentic Reinforcement Learning
Zhengxi Lu, Zhiyuan Yao, Zhuowen Han +8
Reinforcement learning (RL) has emerged as a central paradigm for post-training LLM agents, yet its trajectory-level reward signal provides only coarse supervision for long-horizon…