6 papers
CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers
Weidong Chen, Dexiang Hong, Zhendong Mao +4
The paper introduces CreatiParser, a hybrid generative system that converts raster graphic design images into editable layers (text, background, stickers) using a vision-language m…
Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models
Huatian Zhang, Zhendong Mao, Lei Zhang +1
Direct Preference Optimization (DPO) has proven to be an effective solution for mitigating hallucination in Multimodal Large Language Models (MLLMs) by learning from preference pai…
FACE-net: Factual Calibration and Emotion Augmentation for Retrieval-enhanced Emotional Video Captioning
Weidong Chen, Cheng Ye, Zhendong Mao +5
Emotional Video Captioning (EVC) is an emerging task, which aims to describe factual content with the intrinsic emotions expressed in videos. Existing works perceive global emotion…
LayerEdit: Disentangled Multi-Object Editing via Conflict-Aware Multi-Layer Learning
Fengyi Fu, Mengqi Huang, Lei Zhang +1
Text-driven multi-object image editing which aims to precisely modify multiple objects within an image based on text descriptions, has recently attracted considerable interest. Exi…
DiT: Dynamic Diffusion Transformer for Accurate Image Generation
Weinan Jia, Mengqi Huang, Nan Chen +2
Diffusion models are widely recognized for their ability to generate high-fidelity images. Despite the excellent performance and scalability of the Diffusion Transformer (DiT) arch…
CustomContrast: A Multilevel Contrastive Perspective For Subject-Driven Text-to-Image Customization
Nan Chen, Mengqi Huang, Zhuowei Chen +3
Subject-driven text-to-image (T2I) customization has drawn significant interest in academia and industry. This task enables pre-trained models to generate novel images based on uni…