7 papers
Structure-Semantic Co-optimized Latent Diffusion Model for Fast Visual Anagram Synthesis
Xiang Gao, Yunpeng Jia
The paper introduces a structure-semantic co-optimization framework for latent diffusion models that generates high‑resolution visual anagrams quickly, improving visual quality and…
TEASR: Training-Efficient Any-Step Diffusion Transformer for Real-World Image Super-Resolution
Xiang Gao, Chenxin Zhu, Yushun Fang +2
Diffusion models excel in Real-World Image Super-Resolution (Real-ISR) due to their powerful generative priors but suffer from slow iterative sampling. Although existing one-step d…
FBSDiff++: Improved Frequency Band Substitution of Diffusion Features for Efficient and Highly Controllable Text-Driven Image-to-Image Translation
Xiang Gao, Yunpeng Jia
With large-scale text-to-image (T2I) diffusion models achieving significant advancements in open-domain image creation, increasing attention has been focused on their natural exten…
FBSDiff: Plug-and-Play Frequency Band Substitution of Diffusion Features for Highly Controllable Text-Driven Image Translation
Xiang Gao, Jiaying Liu
Large-scale text-to-image diffusion models have been a revolutionary milestone in the evolution of generative AI and multimodal technology, allowing wonderful image generation with…
LongSplat: Online Generalizable 3D Gaussian Splatting from Long Sequence Images
Guichen Huang, Ruoyu Wang, Xiangjun Gao +4
3D Gaussian Splatting achieves high-fidelity novel view synthesis, but its application to online long-sequence scenarios is still limited. Existing methods either rely on slow per-…
PTDiffusion: Free Lunch for Generating Optical Illusion Hidden Pictures with Phase-Transferred Diffusion Model
Xiang Gao, Shuai Yang, Jiaying Liu
Optical illusion hidden picture is an interesting visual perceptual phenomenon where an image is cleverly integrated into another picture in a way that is not immediately obvious t…