9 papers
SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation
Zikun Qu, Min Zhang, Mingze Kong +5
On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other pla…
Workflow-R1: Group Sub-sequence Policy Optimization for Multi-turn Workflow Construction
Mingze Kong, Zikun Qu, Zhongquan Zhou +7
The rapid evolution of agentic workflows has demonstrated strong performance of LLM-based agents in addressing complex reasoning tasks. However, existing workflow optimization meth…
Dual-Robust Cross-Domain Offline Reinforcement Learning Against Dynamics Shifts
Zhongjian Qiao, Rui Yang, Jiafei Lyu +5
Single-domain offline reinforcement learning (RL) often suffers from limited data coverage, while cross-domain offline RL handles this issue by leveraging additional data from othe…
FedPOB: Sample-Efficient Federated Prompt Optimization via Bandits
Pingchen Lu, Zhi Hong, Zhiwei Shang +6
The performance of large language models (LLMs) is highly sensitive to the input prompt, making prompt optimization a critical task. However, real-world application is hindered by…
Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
Zixuan Huang, Yikun Ban, Lean Fu +4
Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its performance is highly depen…
Convergence Rates of Constrained Expected Improvement
Haowei Wang, Jingyi Wang, Zhongxiang Dai +3
Constrained Bayesian optimization (CBO) methods have seen significant success in black-box optimization with constraints. One of the most commonly used CBO methods is the constrain…