3 papers
cs.AI2026
Rewarding Better Thinking for LLM Preference Alignment
Xubo Liu, Wenya Guo, Ruxue Yan +2
LLM preference alignment aims to optimize models toward human preferences across diverse user instructions. Reinforcement learning has become a major post-training approach for thi…
cs.CV2026
From "What" to "How": Constrained Reasoning for Autoregressive Image Generation
Ruxue Yan, Xubo Liu, Wenya Guo +3
Autoregressive image generation has seen recent improvements with the introduction of chain-of-thought and reinforcement learning. However, current methods merely specify "What" de…
cs.LG2025
ProDS: Preference-oriented Data Selection for Instruction Tuning
Wenya Guo, Zhengkun Zhang, Xumeng Liu +5
Instruction data selection aims to identify a high-quality subset from the training set that matches or exceeds the performance of the full dataset on target tasks. Existing method…