From the 1 of 7 linked papers with an AI index.
7 papers
CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers
Weidong Chen, Dexiang Hong, Zhendong Mao +4
The paper introduces CreatiParser, a hybrid generative system that converts raster graphic design images into editable layers (text, background, stickers) using a vision-language m…
Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models
Huatian Zhang, Zhendong Mao, Lei Zhang +1
Direct Preference Optimization (DPO) has proven to be an effective solution for mitigating hallucination in Multimodal Large Language Models (MLLMs) by learning from preference pai…
FACE-net: Factual Calibration and Emotion Augmentation for Retrieval-enhanced Emotional Video Captioning
Weidong Chen, Cheng Ye, Zhendong Mao +5
Emotional Video Captioning (EVC) is an emerging task, which aims to describe factual content with the intrinsic emotions expressed in videos. Existing works perceive global emotion…
RealCustom++: Representing Images as Real Textual Word for Real-Time Customization
Zhendong Mao, Mengqi Huang, Fei Ding +3
Given a text and an image of a specific subject, text-to-image customization aims to generate new images that align with both the text and the subject's appearance. Existing works…
SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation
Hao Du, Bo Wu, Yan Lu +1
Vision-language temporal alignment is a crucial capability for human dynamic recognition and cognition in real-world scenarios. While existing research focuses on capturing vision-…
CustomContrast: A Multilevel Contrastive Perspective For Subject-Driven Text-to-Image Customization
Nan Chen, Mengqi Huang, Zhuowei Chen +3
Subject-driven text-to-image (T2I) customization has drawn significant interest in academia and industry. This task enables pre-trained models to generate novel images based on uni…