4 papers
PACR: Progressively Ascending Confidence Reward for LLM Reasoning
Eunseop Yoon, Hee Suk Yoon, Jaehyun Jang +5
Reinforcement Learning with Verifiable Rewards (RLVR) has significantly improved LLM reasoning, but its sparse, outcome-based reward provides no guidance for intermediate steps, sl…
Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models
Eunseop Yoon, Hee Suk Yoon, Mark A. Hasegawa-Johnson +1
In the broader context of deep learning, Multimodal Large Language Models have achieved significant breakthroughs by leveraging powerful Large Language Models as a backbone to alig…
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
Hee Suk Yoon, Eunseop Yoon, Mark Hasegawa-Johnson +2
We introduce ConfPO, a method for preference learning in Large Language Models (LLMs) that identifies and optimizes preference-critical tokens based solely on the training policy's…
Physics Informed Distillation for Diffusion Models
Joshua Tian Jin Tee, Kang Zhang, Hee Suk Yoon +3
Diffusion models have recently emerged as a potent tool in generative modeling. However, their inherent iterative nature often results in sluggish image generation due to the requi…