activity
20242026
most citedUnderstanding R1-Zero-Like Training: A Critical Perspective

6 citations · 6 across the 6 of their papers we have counts for

collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2026

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models

Chee Heng Tan, Zhuoyi Lin, Mehul Motani +1

In this paper, we consider the setting where large language models (LLMs) are trained using reinforcement learning (RL) to simultaneously improve reasoning accuracy and verbalize i…

cs.LG2026

Rethinking the Trust Region in LLM Reinforcement Learning

Penghui Qi, Xiangxin Zhou, Zichen Liu +4

Reinforcement learning (RL) has become a cornerstone for fine-tuning Large Language Models (LLMs), with Proximal Policy Optimization (PPO) serving as the de facto standard algorith…

cs.LG2026

Intrinsic Wasserstein Rates for Score-Based Generative Models on Smooth Manifolds

Guoji Fu, Taiji Suzuki, Wee Sun Lee +1

Score-based generative models are trained in high-dimensional ambient spaces, yet many data distributions are supported on low-dimensional nonlinear structures. We prove that, for…

cs.LG2025

Optimizing Anytime Reasoning via Budget Relative Policy Optimization

Penghui Qi, Zichen Liu, Tianyu Pang +3

Scaling test-time compute is crucial for enhancing the reasoning capabilities of large language models (LLMs). Existing approaches typically employ reinforcement learning (RL) to m…

cs.LG2025

Defeating the Training-Inference Mismatch via FP16

Penghui Qi, Zichen Liu, Xiangxin Zhou +4

Reinforcement learning (RL) fine-tuning of large language models (LLMs) often suffers from instability due to the numerical mismatch between the training and inference policies. Wh…

cs.LG2025

Approximation and Generalization Abilities of Score-based Neural Network Generative Models for Sub-Gaussian Distributions

Guoji Fu, Wee Sun Lee

This paper studies the approximation and generalization abilities of score-based neural network generative models (SGMs) in estimating an unknown distribution from i.i.d.…