8 papers
V-Warper: Appearance-Consistent Video Diffusion Personalization via Value Warping
Hyunkoo Lee, Wooseok Jang, Jini Yang +4
Video personalization aims to generate videos that faithfully reflect a user-provided subject while following a text prompt. However, existing approaches often rely on heavy video-…
Data Descriptions from Large Language Models with Influence Estimation
Chaeri Kim, Jaeyeon Bae, Taehwan Kim
Deep learning models have been successful in many areas but understanding their behaviors still remains a black-box. Most prior explainable AI (XAI) approaches have focused on inte…
VEHME: A Vision-Language Model For Evaluating Handwritten Mathematics Expressions
Thu Phuong Nguyen, Duc M. Nguyen, Hyotaek Jeon +4
Automatically assessing handwritten mathematical solutions is an important problem in educational technology with practical applications, but it remains a significant challenge due…
Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
Taesoo Kim, Yongsik Jo, Hyunmin Song +1
Human conversation involves language, speech, and visual cues, with each medium providing complementary information. For instance, speech conveys a vibe or tone not fully captured…
Audio-Guided Visual Editing with Complex Multi-Modal Prompts
Hyeonyu Kim, Seokhoon Jeong, Seonghee Han +2
Visual editing with diffusion models has made significant progress but often struggles with complex scenarios that textual guidance alone could not adequately describe, highlightin…
MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
Junha Song, Yongsik Jo, So Yeon Min +4
Systems such as video chatbots and navigation robots often depend on streaming image captioning to interpret visual inputs. Existing approaches typically employ large multimodal la…