1 citations · 1 across the 11 of their papers we have counts for
4 papers · 1 filter
Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training
Junwon Ko, Dong-Jae Lee, Minchan Kwon +2
LLM agents for sequential decision tasks are often post-trained with trajectory-level outcome labels, but such labels provide little supervision for preserving multiple successful…
</think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination
Seunghee Koh, Sungjae Choi, Minchan Kwon +2
Chain-of-thought (CoT) reasoning improves large reasoning models (LRMs) on complex tasks but often produces long, redundant traces. Recent training-free early-exit methods shorten…
Preference Distillation via Value based Reinforcement Learning
Minchan Kwon, Junwon Ko, Kangil Kim +1
Direct Preference Optimization (DPO) is a powerful paradigm to align language models with human preferences using pairwise comparisons. However, its binary win-or-loss supervision…
StablePrompt: Automatic Prompt Tuning using Reinforcement Learning for Large Language Models
Minchan Kwon, Gaeun Kim, Jongsuk Kim +2
Finding appropriate prompts for the specific task has become an important issue as the usage of Large Language Models (LLM) has expanded. Reinforcement Learning (RL) is widely used…