2 papers
cs.CV2026
Token Bottleneck: One Token to Remember Dynamics
Taekyung Kim, Dongyoon Han, Byeongho Heo +2
Deriving compact and temporally aware visual representations from dynamic scenes is essential for successful execution of sequential scene understanding tasks such as visual tracki…
cs.RO2025
Hierarchical Vision Language Action Model Using Success and Failure Demonstrations
Jeongeun Park, Jihwan Yoon, Byungwoo Jeon +6
Prior Vision-Language-Action (VLA) models are typically trained on teleoperated successful demonstrations, while discarding numerous failed attempts that occur naturally during dat…