2 papers
cs.AI2026
Rewarding Better Thinking for LLM Preference Alignment
Xubo Liu, Wenya Guo, Ruxue Yan +2
LLM preference alignment aims to optimize models toward human preferences across diverse user instructions. Reinforcement learning has become a major post-training approach for thi…
cs.LG2025
ProDS: Preference-oriented Data Selection for Instruction Tuning
Wenya Guo, Zhengkun Zhang, Xumeng Liu +5
Instruction data selection aims to identify a high-quality subset from the training set that matches or exceeds the performance of the full dataset on target tasks. Existing method…