4 papers
THINKSAFE: Self-Generated Safety Alignment for Reasoning Models
Seanie Lee, Sangwoo Park, Yumin Choi +6
Large reasoning models (LRMs) achieve remarkable performance by leveraging reinforcement learning (RL) on reasoning tasks to generate long chain-of-thought (CoT) reasoning. However…
Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models
Sanghyun Lee, Seungryong Kim, Jongho Park +1
Masked Diffusion Models (MDMs) as language models generate by iteratively unmasking tokens, yet their performance crucially depends on the inference time order of unmasking. Prevai…
Effective Test-Time Scaling of Discrete Diffusion through Iterative Refinement
Sanghyun Lee, Sunwoo Kim, Seungryong Kim +2
Test-time scaling through reward-guided generation remains largely unexplored for discrete diffusion models despite its potential as a promising alternative. In this work, we intro…
Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess
Dongyoon Hwang, Hojoon Lee, Jaegul Choo +2
While reinforcement learning (RL) for large language models (LLMs) has shown promise in mathematical reasoning, strategic reasoning for LLMs using RL remains largely unexplored. We…