3 papers
cs.LG2026
Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Road
Ngoc-Hieu Nguyen, Parshin Shojaee, Phuc Minh Nguyen +4
Recent progress in large language models has led to the emergence of reasoning models, which have shown strong performance on complex tasks through specialized fine-tuning procedur…
cs.AI2025
The Reasoning Boundary Paradox: How Reinforcement Learning Constrains Language Models
Phuc Minh Nguyen, Chinh D. La, Duy M. H. Nguyen +3
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a key method for improving Large Language Models' reasoning capabilities, yet recent evidence suggests it may p…
cs.LG2025
Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling
Phuc Minh Nguyen, Ngoc-Hieu Nguyen, Duy H. M. Nguyen +5
Direct Alignment Algorithms (DAAs) such as Direct Preference Optimization (DPO) have emerged as alternatives to the standard Reinforcement Learning from Human Feedback (RLHF) for a…