#diffusion models
115 papers · 1 filter
ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation
Xiao Luo, Mingyang Du, Xin Zhou +5
The paper introduces ROAD, a framework that transfers semantic and structural knowledge from discriminative 3D foundation models into diffusion transformers for 3D shape generation…
S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image
Hail Song, Seokhwan Yang, Jiwon Yang +2
S-Avatar is a method that creates photorealistic 3D head avatars from a single image by using diffusion-guided Gaussian splatting and aligning the result with the FLAME parametric…
TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment
MD Wahiduzzaman Khan, Mingshan Jia, Xiaolin Zhang +2
The paper presents a framework that adds realistic, cross‑identity tongue motion to face reenactment by automatically training a tongue segmentation model and using a spatially con…
FeatFix: Reuse What You Verify through Local Exact-Feature Correction for Faster Cached Diffusion Inference
Hanshuai Cui, Zhiqing Tang, Zhi Yao +3
FeatFix reuses exact intermediate features computed for verification to locally correct draft outputs in cached diffusion inference, speeding up image and video generation while pr…
FDDWAN: A Frequency-Decoupled Diffusion Network for Watermarking Attack
Chunpeng Wang, Yuxin Li, Xiaoyu Wang +3
The paper introduces FDDWAN, a two-stage neural network that removes invisible watermarks by first decomposing images with wavelets and then refining residuals with a diffusion mod…
VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition
Yukun Chen, Tianrui Wang, Zhaoxi Mu +2
The paper introduces VocalRender, a system that can directly synthesize singing voices from musical scores—including lyrics, pitches, note values, and tempo—without needing separat…