Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization
Jingxiao Yang, Wangjie Gan, Yingxuan Zhuang +3
Outcome-based reinforcement learning provides verified feedback for language-model agents, but assigns trajectory-level advantage uniformly to all decisions, yielding coarse credit…
cs.AI2026
AttriMem: Attribution-Guided Process Feedback for Agent Memory Construction
Qinfeng Li, Yuntai Bao, Xinyan Yu +7
Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract, store, update, co…
cs.AI2026
GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification
Wangjie Gan, Miao Pan, Linbo Xi +4
Large language models are typically post-trained using supervised fine-tuning (SFT) and reinforcement learning (RL), yet effectively unifying efficient knowledge injection with rob…