8 papers
SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation
Zikun Qu, Min Zhang, Mingze Kong +5
On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other pla…
Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation
Bo Xue, Zhi Hong, Jiayi Li +3
Large language model (LLM) configuration evaluation is challenging due to limited evaluation budgets, varying costs, and multiple competing objectives. In this paper, we formulate…
Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits
Bo Xue, Ji Cheng, Haodong Jing +2
This paper studies generalized low-rank matrix bandits with multiple prioritized objectives. At each round, the learner selects a matrix-valued arm and observes a vector-valued rew…
DemoPSD: Disagreement-Modulated Policy Self-Distillation
Yunhe Li, Hao Shi, Wenhao Liu +5
On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single model acts as both the teacher and the stud…
T-POP: Test-Time Personalization with Online Preference Feedback
Zikun Qu, Min Zhang, Mingze Kong +7
Personalizing large language models (LLMs) to individual user preferences is a critical step beyond generating generically helpful responses. However, current personalization metho…
Learning to Reason with Insight for Informal Theorem Proving
Yunhe Li, Hao Shi, Bowen Deng +8
Although most of the automated theorem-proving approaches depend on formal proof systems, informal theorem proving can align better with large language models' (LLMs) strength in n…