3 papers
cs.RO2026
MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint Reconstruction
Jung Min Lee, Dohyeok Lee, Seokhun Ju +5
Latent actions learned from diverse human videos serve as pseudo-labels for vision-language-action (VLA) pretraining, but provide effective supervision only if they remain informat…
cs.CV2026
Why Latent Actions Fail, and How to Prevent It
Jung Min Lee, Taehyun Cho, Li Zhao +1
Latent action models (LAMs) aim to learn action-like representations from unlabeled videos by compressing frame-to-frame changes. The frames of in-the-wild videos, however, contain…
cs.RO2025
Learning Generalizable Visuomotor Policy through Dynamics-Alignment
Dohyeok Lee, Jung Min Lee, Munkyung Kim +6
Behavior cloning methods for robot learning suffer from poor generalization due to limited data support beyond expert demonstrations. Recent approaches leveraging video prediction…