2 papers
cs.CL2026
Instruction Tuning with and without Context: Behavioral Shifts and Downstream Impact
Hyunji Lee, Seunghyun Yoon, Yunjae Won +7
Instruction tuning is a widely used approach to improve the instruction-following ability of large language models (LLMs). Instruction-tuning datasets typically include a mixture o…
cs.LG2025
Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization
Yunjae Won, Hyunji Lee, Hyeonbin Hwang +1
Direct Preference Optimization (DPO) has been widely used for aligning language models with human preferences in a supervised manner. However, several key questions remain unresolv…