4 papers
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast
Jiaxin Guo, Yanwei Yue, Xuanbo Fan +2
On-policy self-distillation improves language-model reasoning by querying a teacher on states actually visited by the student. Recent methods create a powerful information asymmetr…
ZeroSearch: Incentivize the Search Capability of LLMs without Searching
Hao Sun, Zile Qiao, Jiayan Guo +7
Effective information searching is essential for enhancing the reasoning and generation capabilities of large language models (LLMs). Recent research has explored using reinforceme…
Mem-T: Densifying Rewards for Long-Horizon Memory Agents
Yanwei Yue, Boci Peng, Xuanbo Fan +3
Memory agents, which depart from predefined memory-processing pipelines by endogenously managing the processing, storage, and retrieval of memories, have garnered increasing attent…
Memory in the Age of AI Agents
Yuyang Hu, Shichun Liu, Yanwei Yue +44
Memory has emerged, and will continue to remain, a core capability of foundation model-based agents. As research on agent memory rapidly expands and attracts unprecedented attentio…