2 papers
cs.LG2025
PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
Mengdi Li, Guanqiao Chen, Xufeng Zhao +3
Reward models (RMs), which are central to existing post-training methods, aim to align LLM outputs with human values by providing feedback signals during fine-tuning. However, exis…
cs.AI2025
Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
Chen Qian, Dongrui Liu, Haochen Wen +3
Large reasoning models (LRMs) have demonstrated impressive capabilities in complex problem-solving, yet their internal reasoning mechanisms remain poorly understood. In this paper,…