3 papers
cs.LG2026
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
Gengsheng Li, Tianyu Yang, Junfeng Fang +6
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models. While Group Relative Policy Optimization (GRPO) is wid…
cs.LG2026
NExT-Guard: Training-Free Streaming Safeguard without Token-Level Labels
Junfeng Fang, Nachuan Chen, Houcheng Jiang +5
Large language models are increasingly deployed in streaming scenarios, rendering conventional post-hoc safeguards ineffective as they fail to interdict unsafe content in real-time…
cs.LG2025
Generative Uncertainty in Diffusion Models
Metod Jazbec, Eliot Wong-Toi, Guoxuan Xia +3
Diffusion models have recently driven significant breakthroughs in generative modeling. While state-of-the-art models produce high-quality samples on average, individual samples ca…