3 papers
cs.AI2025
PACR: Progressively Ascending Confidence Reward for LLM Reasoning
Eunseop Yoon, Hee Suk Yoon, Jaehyun Jang +5
Reinforcement Learning with Verifiable Rewards (RLVR) has significantly improved LLM reasoning, but its sparse, outcome-based reward provides no guidance for intermediate steps, sl…
cs.CV2025
Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models
Eunseop Yoon, Hee Suk Yoon, Mark A. Hasegawa-Johnson +1
In the broader context of deep learning, Multimodal Large Language Models have achieved significant breakthroughs by leveraging powerful Large Language Models as a backbone to alig…
cs.CL2025
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
Hee Suk Yoon, Eunseop Yoon, Mark Hasegawa-Johnson +2
We introduce ConfPO, a method for preference learning in Large Language Models (LLMs) that identifies and optimizes preference-critical tokens based solely on the training policy's…