7 papers
Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
Hieu Trung Nguyen, Bao Nguyen, Wenao Ma +3
Sampling efficiency is a key bottleneck in reinforcement learning with verifiable rewards. Existing group-based policy optimization methods, such as GRPO, allocate a fixed number o…
Reasoning Planning for Language Models
Bao Nguyen, Hieu Trung Nguyen, Ruifeng She +2
Selecting an appropriate reasoning method for a given query remains a key challenge in language model generation. Existing approaches typically generate multiple candidate response…
Distributional Surgery for Language Model Activations
Bao Nguyen, Binh Nguyen, Duy Nguyen +1
Language models, while capable of generating remarkably coherent and seemingly accurate text, can occasionally produce undesirable content, including harmful or toxic outputs. In t…
Structured Pruning for Diverse Best-of-N Reasoning Optimization
Hieu Trung Nguyen, Bao Nguyen, Viet Anh Nguyen
Model pruning in transformer-based language models, traditionally viewed as a means of achieving computational savings, can enhance the model's reasoning capabilities. In this work…
Task-driven Layerwise Additive Activation Intervention
Hieu Trung Nguyen, Bao Nguyen, Binh Nguyen +1
Modern language models (LMs) have significantly advanced generative modeling in natural language processing (NLP). Despite their success, LMs often struggle with adaptation to new…
Probe-Free Low-Rank Activation Intervention
Chonghe Jiang, Bao Nguyen, Anthony Man-Cho So +1
Language models (LMs) can produce texts that appear accurate and coherent but contain untruthful or toxic content. Inference-time interventions that edit the hidden activations hav…