Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Preference Distillation via Value based Reinforcement Learning
Minchan Kwon, Junwon Ko, Kangil Kim +1
Direct Preference Optimization (DPO) is a powerful paradigm to align language models with human preferences using pairwise comparisons. However, its binary win-or-loss supervision…
cs.CL2024
StablePrompt: Automatic Prompt Tuning using Reinforcement Learning for Large Language Models
Minchan Kwon, Gaeun Kim, Jongsuk Kim +2
Finding appropriate prompts for the specific task has become an important issue as the usage of Large Language Models (LLM) has expanded. Reinforcement Learning (RL) is widely used…