7 papers
Retrieve, Don't Retrain: Extending Vision Language Action Models to New Tasks at Test Time
Jeongeun Park, Juhan Park, Taekyung Kim +3
Extending a vision-language-action (VLA) policy to a new task typically requires task-specific teleoperated demonstrations and per-task fine-tuning, making adaptation costly in bot…
Exploring Conditions for Diffusion models in Robotic Control
Heeseong Shin, Byeongho Heo, Dongyoon Han +2
While pre-trained visual representations have significantly advanced imitation learning, they are often task-agnostic as they remain frozen during policy learning. In this work, we…
MuCo: Multi-turn Contrastive Learning for Multimodal Embedding Model
Geonmo Gu, Byeongho Heo, Jaemyung Yu +7
Universal Multimodal embedding models built on Multimodal Large Language Models (MLLMs) have traditionally employed contrastive learning, which aligns representations of query-targ…
VisualScratchpad: Inference-time Visual Concepts Analysis in Vision Language Models
Hyesu Lim, Jinho Choi, Taekyung Kim +3
High-performing vision language models still produce incorrect answers, yet their failure modes are often difficult to explain. To make model internals more accessible and enable s…
Is Your Safe Controller Actually Safe? A Critical Review of CBF Tautologies and Hidden Assumptions
Taekyung Kim
This tutorial provides a critical review of the practical application of Control Barrier Functions (CBFs) in robotic safety. While the theoretical foundations of CBFs are well-esta…
Token Bottleneck: One Token to Remember Dynamics
Taekyung Kim, Dongyoon Han, Byeongho Heo +2
Deriving compact and temporally aware visual representations from dynamic scenes is essential for successful execution of sequential scene understanding tasks such as visual tracki…