collaborators

7 papers

cs.AI2026

A Regret Minimization Framework on Preference Learning in Large Language Models

Suhwan Kim, Taehyun Cho, Geon-Hyeong Kim +4

Reinforcement learning with verifiable rewards (RLVR) has enabled progress on reasoning-intensive tasks by relying on task-specific verifiers that provide automated correctness sig…

cs.RO2026

MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint Reconstruction

Jung Min Lee, Dohyeok Lee, Seokhun Ju +5

Latent actions learned from diverse human videos serve as pseudo-labels for vision-language-action (VLA) pretraining, but provide effective supervision only if they remain informat…

cs.LG2026

Probabilistic Smoothing with Ratio-Monotone Transforms for Global Optimization

Kukyoung Jang, Taehyun Cho, Junrui Zhang +2

Probabilistic smoothing is a standard tool for global optimization, but existing methods rely on Gaussian kernels and specific transforms, often resulting in strong hyperparameter…

cs.CV2026

Why Latent Actions Fail, and How to Prevent It

Jung Min Lee, Taehyun Cho, Li Zhao +1

Latent action models (LAMs) aim to learn action-like representations from unlabeled videos by compressing frame-to-frame changes. The frames of in-the-wild videos, however, contain…

cs.RO2025

Learning Generalizable Visuomotor Policy through Dynamics-Alignment

Dohyeok Lee, Jung Min Lee, Munkyung Kim +6

Behavior cloning methods for robot learning suffer from poor generalization due to limited data support beyond expert demonstrations. Recent approaches leveraging video prediction…

cs.LG2025

Policy-labeled Preference Learning: Is Preference Enough for RLHF?

Taehyun Cho, Seokhun Ju, Seungyub Han +3

To design rewards that align with human goals, Reinforcement Learning from Human Feedback (RLHF) has emerged as a prominent technique for learning reward functions from human prefe…