5 papers · 1 filter
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
Yuxiao Yang, Tianrun Yu, Shangzhe Li +6
We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify \emph{termination…
Provable and Practical In-Context Policy Optimization for Self-Improvement
Tianrun Yu, Yuxiao Yang, Zhaoyang Wang +6
We study test-time scaling, where a model improves its answer through multi-round self-reflection at inference. We introduce In-Context Policy Optimization (ICPO), in which an agen…
SynthAgent: Adapting Web Agents with Synthetic Supervision
Zhaoyang Wang, Yiming Liang, Xuchao Zhang +9
Web agents struggle to adapt to new websites due to the scarcity of environment specific tasks and demonstrations. Recent works have explored synthetic data generation to address t…
Anyprefer: An Agentic Framework for Preference Data Synthesis
Yiyang Zhou, Zhaoyang Wang, Tianle Wang +13
High-quality preference data is essential for aligning foundation models with human values through preference learning. However, manual annotation of such data is often time-consum…
CREAM: Consistency Regularized Self-Rewarding Language Models
Zhaoyang Wang, Weilei He, Zhiyuan Liang +5
Recent self-rewarding large language models (LLM) have successfully applied LLM-as-a-Judge to iteratively improve the alignment performance without the need of human annotations fo…