activity
20242026
collaborators

5 papers

cs.CV2026

Disco-LoRA: Disentangled Composition of Content, Style, and Motion for Multi-concept Video Customization

Xuancheng Xu, Gengyun Jia, Bing-Kun Bao

Video customization based on Text-to-Video (T2V) models aims to learn specific features from reference data to generate controllable videos. While significant strides have been mad…

cs.CV2025

SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation

Xuancheng Xu, Yaning Li, Sisi You +1

Customized video generation aims to produce videos that faithfully preserve the subject's appearance from reference images while maintaining temporally consistent motion from refer…

cs.CV2025

Chain-of-Cooking:Cooking Process Visualization via Bidirectional Chain-of-Thought Guidance

Mengling Xu, Ming Tao, Bing-Kun Bao

Cooking process visualization is a promising task in the intersection of image generation and food analysis, which aims to generate an image for each cooking step of a recipe. Howe…

cs.CV2024

Do We Need to Design Specific Diffusion Models for Different Tasks? Try ONE-PIC

Ming Tao, Bing-Kun Bao, Yaowei Wang +1

Large pretrained diffusion models have demonstrated impressive generation capabilities and have been adapted to various downstream tasks. However, unlike Large Language Models (LLM…

cs.CV2024

StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion

Ming Tao, Bing-Kun Bao, Hao Tang +2

Story visualization aims to generate a series of realistic and coherent images based on a storyline. Current models adopt a frame-by-frame architecture by transforming the pre-trai…