5 papers
Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention
Yang Yu, Zhuangzhuang Chen, Lanqing Li +1
Recently, reinforcement learning (RL) has become a common choice in enhancing the reasoning capabilities of vision-language models (VLMs). Considering existing RL-based finetuning…
Into the Gray Zone: Domain Contexts Can Blur LLM Safety Boundaries
Ki Sen Hung, Xi Yang, Chang Liu +7
A central goal of LLM alignment is to balance helpfulness with harmlessness, yet these objectives conflict when the same knowledge serves both legitimate and malicious purposes. Th…
Parallelism and Generation Order in Masked Diffusion Language Models: Limits Today, Potential Tomorrow
Yangyang Zhong, Yanmei Gu, Zhengqing Zang +14
Masked Diffusion Language Models (MDLMs) promise parallel token generation and arbitrary-order decoding, yet it remains unclear to what extent current models truly realize these ca…
Distribution-Aware Reward Estimation for Test-Time Reinforcement Learning
Bodong Du, Xuanqi Huang, Xiaomeng Li
Test-time reinforcement learning (TTRL) enables large language models (LLMs) to self-improve on unlabeled inputs, but its effectiveness critically depends on how reward signals are…
Scaling and Transferability of Annealing Strategies in Large Language Model Training
Siqi Wang, Zhengyu Chen, Teng Xiao +5
Learning rate scheduling is crucial for training large language models, yet understanding the optimal annealing strategies across different model configurations remains challenging…