10 papers
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment
Xiuyu Li, Jinkai Zhang, Mingyang Yi +4
Reinforcement Learning (RL) post-training alignment for language models is effective, but also costly and unstable in practice, owing to its complicated training process. To addres…
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
Lai Wei, Liangbo He, Jun Lan +9
Multimodal Large Language Models (MLLMs) excel at broad visual understanding but still struggle with fine-grained perception, where decisive evidence is small and easily overwhelme…
Stabilizing Policy Gradient Methods via Reward Profiling
Shihab Ahmed, El Houcine Bergou, Aritra Dutta +1
Policy gradient methods, which have been extensively studied in the last decade, offer an effective and efficient framework for reinforcement learning problems. However, their perf…
Towards a Theoretical Understanding to the Generalization of RLHF
Zhaochun Li, Mingyang Yi, Yue Wang +2
Reinforcement Learning from Human Feedback (RLHF) and its variants have emerged as the dominant approaches for aligning Large Language Models with human intent. While empirically e…
Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
Xuerui Su, Shufang Xie, Guoqing Liu +7
Recently, Large Language Models (LLMs) have rapidly evolved, approaching Artificial General Intelligence (AGI) while benefiting from large-scale reinforcement learning to enhance H…
Pre-training Generative Recommender with Multi-Identifier Item Tokenization
Bowen Zheng, Enze Liu, Zhongfu Chen +4
Generative recommendation autoregressively generates item identifiers to recommend potential items. Existing methods typically adopt a one-to-one mapping strategy, where each item…