1 citations · 1 across the 10 of their papers we have counts for
15 papers
Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
Hieu Trung Nguyen, Bao Nguyen, Wenao Ma +3
Sampling efficiency is a key bottleneck in reinforcement learning with verifiable rewards. Existing group-based policy optimization methods, such as GRPO, allocate a fixed number o…
Exploring Diverse Generation Paths via Inference-time Stiefel Activation Steering
Dongxuan Zhu, Ly Tran Ho Khanh, Andy Yat-Ming Cheung +2
Language models often default to a narrow set of high-probability outputs, leaving their generation paths homogeneous and prone to mode collapse. Sampling-based strategies inject r…
SCOPE: Spectral Concentration by Distributionally Robust Joint Covariance-Precision Estimation
Renjie Chen, Viet Anh Nguyen, Huifu Xu
We propose a distributionally robust formulation for simultaneously estimating the covariance matrix and the precision matrix of a random vector.The proposed model minimizes the wo…
Test-time Diverse Reasoning by Riemannian Activation Steering
Ly Tran Ho Khanh, Dongxuan Zhu, Man-Chung Yue +1
Best-of- reasoning improves the accuracy of language models in solving complex tasks by sampling multiple candidate solutions and then selecting the best one based on some crite…
Reasoning Planning for Language Models
Bao Nguyen, Hieu Trung Nguyen, Ruifeng She +2
Selecting an appropriate reasoning method for a given query remains a key challenge in language model generation. Existing approaches typically generate multiple candidate response…
Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
Quan Dao, Xiaoxiao He, Ligong Han +6
Visual autoregressive models (VAR) have recently emerged as a promising class of generative models, achieving performance comparable to diffusion models in text-to-image generation…