7 papers
Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models
Sangoh Lee, Sangwoo Mo, Wook-Shin Han
Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which m…
Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting
Sangoh Lee, Sangwoo Mo, Wook-Shin Han
While Vision-Language-Action (VLA) models generalize well to generic instructions, they struggle with personalized commands such as "bring my cup," where the robot must act on one…
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
Byeongju Woo, Zilin Wang, Byeonghyun Pak +2
Vision-language models such as CLIP often struggle to faithfully understand long, detail-rich captions, relying on dominant scene cues while overlooking fine-grained visual evidenc…
Topology-Aware Representation Alignment for Semi-Supervised Vision-Language Learning
Junwon You, Mihyun Jang, Sangwoo Mo +1
Vision-language models have shown strong performance, but they often generalize poorly to specialized domains. While semi-supervised vision-language learning mitigates this limitat…
Multimodal Dataset Distillation Made Simple by Prototype-Guided Data Synthesis
Junhyeok Choi, Sangwoo Mo, Minwoo Chae
Recent advances in multimodal learning have achieved remarkable success across diverse vision-language tasks. However, such progress heavily relies on large-scale image-text datase…
SHED Light on Segmentation for Dense Prediction
Seung Hyun Lee, Sangwoo Mo, Stella X. Yu
Dense prediction infers per-pixel values from a single image and is fundamental to 3D perception and robotics. Although real-world scenes exhibit strong structure, existing methods…