2 citations · 3 across the 18 of their papers we have counts for
7 papers · 1 filter
DramaDirector: Geometry-Guided Short Drama Generation
Hengji Zhou, Sijie Liu, Jianrun Chen +3
Short dramas, with their rapid shot rhythms, dialogue-driven focus shifts, and demanding cinematographic grounding, pose challenges that prompt-level or text-only video generation…
Navigating User Behavior toward Personalized Multimodal Generation
Hengji Zhou, Yufeng Liu, Ye Liu +3
Modern AIGC pipelines deliver high-fidelity images and videos but presuppose a well-formed creation instruction, while end users rarely articulate visual details, leaving generator…
TailorMind: Towards Preference-Aligned Multimodal Content Generation
Hengji Zhou, Ye Liu, Yufeng Liu +3
Personalized content systems depend on available UGC and struggle when suitable content is absent, delayed, or costly to create. Although multimodal generators can synthesize conte…
TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval
Zixu Li, Yupeng Hu, Zhiheng Fu +3
Composed Image Retrieval (CIR) is an important image retrieval paradigm that enables users to retrieve a target image using a multimodal query that consists of a reference image an…
SJD-VP: Speculative Jacobi Decoding with Verification Prediction for Autoregressive Image Generation
Bingqi Shan, Baoquan Zhang, Xiaochen Qi +3
Speculative Jacobi Decoding (SJD) has emerged as a promising method for accelerating autoregressive image generation. Despite its potential, existing SJD approaches often suffer fr…
AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation
Dongjie Cheng, Ruifeng Yuan, Yongqi Li +5
Real-world perception and interaction are inherently multimodal, encompassing not only language but also vision and speech, which motivates the development of "Omni" MLLMs that sup…