8 papers
Back to the Familiar Future: Failure Recovery for VLA Policies via Pre-Imagined Milestone Selection
Suyeon Shin, Juwon Kim, Hyeonbin Park +4
Vision-language-action (VLA) policies can deviate from nominal trajectories during manipulation, even when tasks remain physically feasible. Recovering from these deviations is cha…
Surface-Based Visibility-Guided Uncertainty for Continuous Active 3D Neural Reconstruction
Hyunseo Kim, Hyeonseo Yang, Taekyung Kim +4
View selection is critical in active 3D neural reconstruction as it impacts the contents of training set and resulting final output quality. Recent view selection strategies emphas…
Continual Vision-and-Language Navigation
Seongjun Jeong, Gi-Cheon Kang, Seongho Choi +2
Developing Vision-and-Language Navigation (VLN) agents typically assumes a \textit{train-once-deploy-once} strategy, which is unrealistic as deployed agents continually encounter n…
OCK: Unsupervised Dynamic Video Prediction with Object-Centric Kinematics
Yeon-Ji Song, Jaein Kim, Suhyung Choi +2
Human perception involves decomposing complex multi-object scenes into time-static object appearance (i.e., size, shape, color) and time-varying object motion (i.e., position, velo…
DBMovi-GS: Dynamic View Synthesis from Blurry Monocular Video via Sparse-Controlled Gaussian Splatting
Yeon-Ji Song, Jaein Kim, Byung-Ju Kim +1
Novel view synthesis is a task of generating scenes from unseen perspectives; however, synthesizing dynamic scenes from blurry monocular videos remains an unresolved challenge that…
CLIP-RT: Learning Language-Conditioned Robotic Policies from Natural Language Supervision
Gi-Cheon Kang, Junghyun Kim, Kyuhwan Shim +2
Teaching robots desired skills in real-world environments remains challenging, especially for non-experts. A key bottleneck is that collecting robotic data often requires expertise…