8 papers
InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion
Hoiyeong Jin, Hyojin Jang, Junha Hyung +6
Recent advances in diffusion models have enabled impressive video editing capabilities, yet production-grade Video Object Insertion (VOI) remains challenging due to inadequate 4D s…
Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement
Kinam Kim, Namiko Saito, Heecheol Kim +3
Vision-Language-Action (VLA) models can generalize across diverse manipulation tasks, but their imitation-learning-based policies remain brittle in precise physical interactions du…
FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control
Donghu Kim, Youngdo Lee, Minho Park +10
Reinforcement learning (RL) is a core approach for robot control when expert demonstrations are unavailable. On-policy methods such as Proximal Policy Optimization (PPO) are widely…
ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models
Minho Park, Kinam Kim, Junha Hyung +5
Diffusion and flow matching models have emerged as powerful robot policies, enabling Vision-Language-Action (VLA) models to generalize across diverse scenes and instructions. Yet,…
Temporal In-Context Fine-Tuning with Temporal Reasoning for Versatile Control of Video Diffusion Models
Kinam Kim, Junha Hyung, Jaegul Choo
Recent advances in text-to-video diffusion models have enabled high-quality video synthesis, but controllable generation remains challenging, particularly under limited data and co…
EgoX: Egocentric Video Generation from a Single Exocentric Video
Taewoong Kang, Kinam Kim, Dohyeon Kim +3
Egocentric perception enables humans to experience and understand the world directly from their own point of view. Translating exocentric (third-person) videos into egocentric (fir…