6 papers
CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image Editing
Yan Li, Lin Liu, Xiaopeng Zhang +4
Instruction-based image editing with diffusion models has achieved impressive results, yet existing methods struggle with fine-grained instructions specifying precise attributes su…
Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space
Yan Li, Changyao Tian, Renqiu Xia +7
We propose AdapTok, an adaptive temporal causal video tokenizer that can flexibly allocate tokens for different frames based on video content. AdapTok is equipped with a block-wise…
ESDiff: Encoding Strategy-inspired Diffusion Model with Few-shot Learning for Color Image Inpainting
Junyan Zhang, Yan Li, Mengxiao Geng +2
Image inpainting is a technique used to restore missing or damaged regions of an image. Traditional methods primarily utilize information from adjacent pixels for reconstructing mi…
Half-order Fine-Tuning for Diffusion Model: A Recursive Likelihood Ratio Optimizer
Tao Ren, Zishi Zhang, Jingyang Jiang +9
The probabilistic diffusion model (DM), generating content by inferencing through a recursive chain structure, has emerged as a powerful framework for visual generation. After pre-…
SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model
Yan Li, Ziya Zhou, Zhiqiang Wang +3
Recent advancements in generative models have significantly enhanced talking face video generation, yet singing video generation remains underexplored. The differences between huma…
HiPrompt: Tuning-free Higher-Resolution Generation with Hierarchical MLLM Prompts
Xinyu Liu, Yingqing He, Lanqing Guo +10
The potential for higher-resolution image generation using pretrained diffusion models is immense, yet these models often struggle with issues of object repetition and structural a…