2 papers
cs.CV2025
Multi-dimensional Preference Alignment by Conditioning Reward Itself
Jiho Jang, Jinyoung Kim, Kyungjune Baek +1
Reinforcement Learning from Human Feedback has emerged as a standard for aligning diffusion models. However, we identify a fundamental limitation in the standard DPO formulation be…
cs.LG2025
Understanding Differential Transformer Unchains Pretrained Self-Attentions
Chaerin Kong, Jiho Jang, Nojun Kwak
Differential Transformer has recently gained significant attention for its impressive empirical performance, often attributed to its ability to perform noise canceled attention. Ho…