3 papers
cs.AI2026
A Regret Minimization Framework on Preference Learning in Large Language Models
Suhwan Kim, Taehyun Cho, Geon-Hyeong Kim +4
Reinforcement learning with verifiable rewards (RLVR) has enabled progress on reasoning-intensive tasks by relying on task-specific verifiers that provide automated correctness sig…
cs.LG2025
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
Geon-Hyeong Kim, Yu Jin Kim, Byoungjip Kim +4
As Large Language Models (LLMs) are increasingly deployed in real-world applications, balancing helpfulness and safety has become a central challenge. A natural approach is to inco…
cs.AI2024
Semantic Skill Grounding for Embodied Instruction-Following in Cross-Domain Environments
Sangwoo Shin, Seunghyun Kim, Youngsoo Jang +2
In embodied instruction-following (EIF), the integration of pretrained language models (LMs) as task planners emerges as a significant branch, where tasks are planned at the skill…