4 papers
Hierarchical Vision Language Action Model Using Success and Failure Demonstrations
Jeongeun Park, Jihwan Yoon, Byungwoo Jeon +6
Prior Vision-Language-Action (VLA) models are typically trained on teleoperated successful demonstrations, while discarding numerous failed attempts that occur naturally during dat…
ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
Huiwon Jang, Sihyun Yu, Heeseung Kwon +3
Leveraging temporal context is crucial for success in partially observable robotic tasks. However, prior work in behavior cloning has demonstrated inconsistent performance gains wh…
Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics
Dongyoung Kim, Sumin Park, Huiwon Jang +3
Large Vision-Language Models (LVLMs) have recently shown great promise in advancing robotics by combining embodied reasoning with robot control. A common approach involves training…
Learning Multi-frame and Monocular Prior for Estimating Geometry in Dynamic Scenes
Seong Hyeon Park, Jinwoo Shin
In monocular videos that capture dynamic scenes, estimating the 3D geometry of video contents has been a fundamental challenge in computer vision. Specifically, the task is signifi…