7 papers
3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance
Dongyoon Hwang, Byungkun Lee, Dongjin Kim +7
Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation. Recent work in this paradigm u…
InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion
Hoiyeong Jin, Hyojin Jang, Junha Hyung +6
Recent advances in diffusion models have enabled impressive video editing capabilities, yet production-grade Video Object Insertion (VOI) remains challenging due to inadequate 4D s…
AHS: Adaptive Head Synthesis via Synthetic Data Augmentations
Taewoong Kang, Hyojin Jang, Sohyun Jeong +4
Recent digital media advancements have created increasing demands for sophisticated portrait manipulation techniques, particularly head swapping, where one's head is seamlessly int…
ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models
Minho Park, Kinam Kim, Junha Hyung +5
Diffusion and flow matching models have emerged as powerful robot policies, enabling Vision-Language-Action (VLA) models to generalize across diverse scenes and instructions. Yet,…
Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models
Sungwon Hwang, Hyojin Jang, Kinam Kim +2
Fine-tuning Video Diffusion Models (VDMs) at the user level to generate videos that reflect specific attributes of training data presents notable challenges, yet remains underexplo…
SurFhead: Affine Rig Blending for Geometrically Accurate 2D Gaussian Surfel Head Avatars
Jaeseong Lee, Taewoong Kang, Marcel C. Bühler +5
Recent advancements in head avatar rendering using Gaussian primitives have achieved significantly high-fidelity results. Although precise head geometry is crucial for applications…