From the 1 of 7 linked papers with an AI index.
7 papers
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
Yifan Ding, Xincheng Wei, Yoshua Y. Li +7
Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stron…
FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification
Haoqing Wang, Xingrun Xing, Wei Xia +2
FaithEyes proposes a multi‑agent framework where a vision‑language model judges its own tool calls to ensure they are useful, improving both accuracy and tool faithfulness on visua…
Trust Region On-Policy Distillation
Xingrun Xing, Haoqing Wang, Boyan Gao +2
On-Policy Distillation (OPD) is a fundamental technique for efficient post-training of large language models (LLMs), with broad applications in agent learning, multi-task enhanceme…
Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning
Ziheng Li, Liu Kang, Feng Xiao +7
Group Relative Policy Optimization (GRPO) has emerged as a promising critic-free reinforcement learning paradigm for reasoning tasks. However, standard GRPO employs a coarse-graine…
MemTrain: Self-Supervised Context Memory Training
Ziheng Li, Xingrun Xing, Haoqing Wang +2
Memory is an indispensable capability for long-horizon LLM agents, enabling them to preserve and utilize information accumulated across extended interactions. Existing memory-agent…
IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models
Haonan Song, Qingchen Xie, Huan Zhu +12
Generative Reward Models (GRMs) have demonstrated strong performance in reward modeling, due to their interpretability and potential for refinement through reinforcement learning (…