activity
20242026
collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

PixelU: A U-Shaped Transformer for Efficient End-to-End Pixel Diffusion

Zipeng Guo, Lichen Ma, Yu He +4

End-to-end pixel-space diffusion models bypass the lossy compression of Latent Diffusion Models (LDMs) but struggle to jointly model low-frequency semantics and high-frequency sign…

cs.CV2025

CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image Editing

Yan Li, Lin Liu, Xiaopeng Zhang +4

Instruction-based image editing with diffusion models has achieved impressive results, yet existing methods struggle with fine-grained instructions specifying precise attributes su…

cs.CV2025

Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space

Yan Li, Changyao Tian, Renqiu Xia +7

We propose AdapTok, an adaptive temporal causal video tokenizer that can flexibly allocate tokens for different frames based on video content. AdapTok is equipped with a block-wise…

cs.CV2025

ESDiff: Encoding Strategy-inspired Diffusion Model with Few-shot Learning for Color Image Inpainting

Junyan Zhang, Yan Li, Mengxiao Geng +2

Image inpainting is a technique used to restore missing or damaged regions of an image. Traditional methods primarily utilize information from adjacent pixels for reconstructing mi…

cs.CV2025

Half-order Fine-Tuning for Diffusion Model: A Recursive Likelihood Ratio Optimizer

Tao Ren, Zishi Zhang, Jingyang Jiang +9

The probabilistic diffusion model (DM), generating content by inferencing through a recursive chain structure, has emerged as a powerful framework for visual generation. After pre-…

cs.CV2024

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model

Yan Li, Ziya Zhou, Zhiqiang Wang +3

Recent advancements in generative models have significantly enhanced talking face video generation, yet singing video generation remains underexplored. The differences between huma…