10 papers
Environmental Understanding Vision-Language Model for Embodied Agent
Jinsik Bang, Jaeyeon Bae, Donggyu Lee +2
Vision-language models (VLMs) have shown strong perception and reasoning abilities for instruction-following embodied agents. However, despite these abilities and their generalizat…
Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video
Chanhyuk Choi, Taesoo Kim, Donggyu Lee +2
Talking face generation has gained significant attention as a core application of generative models. To enhance the expressiveness and realism of synthesized videos, emotion editin…
V-Warper: Appearance-Consistent Video Diffusion Personalization via Value Warping
Hyunkoo Lee, Wooseok Jang, Jini Yang +4
Video personalization aims to generate videos that faithfully reflect a user-provided subject while following a text prompt. However, existing approaches often rely on heavy video-…
MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
Junha Song, Yongsik Jo, So Yeon Min +4
Systems such as video chatbots and navigation robots often depend on streaming image captioning to interpret visual inputs. Existing approaches typically employ large multimodal la…
Data Descriptions from Large Language Models with Influence Estimation
Chaeri Kim, Jaeyeon Bae, Taehwan Kim
Deep learning models have been successful in many areas but understanding their behaviors still remains a black-box. Most prior explainable AI (XAI) approaches have focused on inte…
VEHME: A Vision-Language Model For Evaluating Handwritten Mathematics Expressions
Thu Phuong Nguyen, Duc M. Nguyen, Hyotaek Jeon +4
Automatically assessing handwritten mathematical solutions is an important problem in educational technology with practical applications, but it remains a significant challenge due…