4 papers
Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models
Zhaoyi Li, Deyang Kong, Yuan Wei +13
On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly unders…
Hybrid Open-Ended Tri-Evolution Makes Better Deep Researcher
Hongming Piao, Chi Liu, Mengzhuo Chen +5
Deep research and agent evolution serve as de-facto tasks for AI agents in real-world applications toward artificial general intelligence. The former enables autonomous retrieval a…
Anyprefer: An Agentic Framework for Preference Data Synthesis
Yiyang Zhou, Zhaoyang Wang, Tianle Wang +13
High-quality preference data is essential for aligning foundation models with human values through preference learning. However, manual annotation of such data is often time-consum…
CREAM: Consistency Regularized Self-Rewarding Language Models
Zhaoyang Wang, Weilei He, Zhiyuan Liang +5
Recent self-rewarding large language models (LLM) have successfully applied LLM-as-a-Judge to iteratively improve the alignment performance without the need of human annotations fo…