3 papers
cs.CV2025
ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
Runhui Huang, Chunwei Wang, Junwei Yang +8
We present ILLUME+ that leverages dual visual tokenization and a diffusion decoder to improve both deep semantic understanding and high-fidelity image generation. Existing unified…
eess.IV2025
FreqPrior: Improving Video Diffusion Models with Frequency Filtering Gaussian Noise
Yunlong Yuan, Yuanfan Guo, Chunwei Wang +3
Text-driven video generation has advanced significantly due to developments in diffusion models. Beyond the training and sampling phases, recent studies have investigated noise pri…
cs.CV2025
Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising
Yunlong Yuan, Yuanfan Guo, Chunwei Wang +2
Recent advances in diffusion models have greatly improved text-driven video generation. However, training models for long video generation demands significant computational power a…