activity
20242026
collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

Yuxiao Yang, Tianrun Yu, Shangzhe Li +6

We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify \emph{termination…

cs.LG2026

Provable and Practical In-Context Policy Optimization for Self-Improvement

Tianrun Yu, Yuxiao Yang, Zhaoyang Wang +6

We study test-time scaling, where a model improves its answer through multi-round self-reflection at inference. We introduce In-Context Policy Optimization (ICPO), in which an agen…

cs.LG2025

SynthAgent: Adapting Web Agents with Synthetic Supervision

Zhaoyang Wang, Yiming Liang, Xuchao Zhang +9

Web agents struggle to adapt to new websites due to the scarcity of environment specific tasks and demonstrations. Recent works have explored synthetic data generation to address t…

cs.LG2025

Anyprefer: An Agentic Framework for Preference Data Synthesis

Yiyang Zhou, Zhaoyang Wang, Tianle Wang +13

High-quality preference data is essential for aligning foundation models with human values through preference learning. However, manual annotation of such data is often time-consum…

cs.LG2024

CREAM: Consistency Regularized Self-Rewarding Language Models

Zhaoyang Wang, Weilei He, Zhiyuan Liang +5

Recent self-rewarding large language models (LLM) have successfully applied LLM-as-a-Judge to iteratively improve the alignment performance without the need of human annotations fo…