4 papers
A Tilted Seesaw: Revisiting Autoencoder Trade-off for Controllable Diffusion
Pu Cao, Yiyang Ma, Feng Zhou +3
In latent diffusion models, the autoencoder (AE) is typically expected to balance two capabilities: faithful reconstruction and a generation-friendly latent space (e.g., low gFID).…
ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers
Yiyang Ma, Feng Zhou, Xuedan Yin +3
Leveraging pre-trained Diffusion Transformers (DiTs) for high-resolution (HR) image synthesis often leads to spatial layout collapse and degraded texture fidelity. Prior work mitig…
Light-SQ: Structure-aware Shape Abstraction with Superquadrics for Generated Meshes
Yuhan Wang, Weikai Chen, Zeyu Hu +19
In user-generated-content (UGC) applications, non-expert users often rely on image-to-3D generative models to create 3D assets. In this context, primitive-based shape abstraction o…
Exploring Position Encoding in Diffusion U-Net for Training-free High-resolution Image Generation
Feng Zhou, Pu Cao, Yiyang Ma +2
Denoising higher-resolution latents via a pre-trained U-Net leads to repetitive and disordered image patterns. Although recent studies make efforts to improve generative quality by…