7 papers
MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation
Zhiyuan Xia, Haojie Li, Jingyu Lin +2
Thermal-to-visible face translation presents fundamental challenges including geometric discontinuities, semantic attribute mismatches, and identity degradation. We propose MTVDiff…
HERO: Hierarchical Extrapolation and Refresh for Efficient World Models
Quanjian Song, Xinyu Wang, Donghao Zhou +3
Generation-driven world models create immersive virtual environments but suffer slow inference due to the iterative nature of diffusion models. While recent advances have improved…
IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation
Donghao Zhou, Jingyu Lin, Guibao Shen +8
Recent visual generative models enable story generation with consistent characters from text, but human-centric story generation faces additional challenges, such as maintaining de…
A Gray-box Attack against Latent Diffusion Model-based Image Editing by Posterior Collapse
Zhongliang Guo, Chun Tong Lei, Lei Fang +7
Recent advancements in Latent Diffusion Models (LDMs) have revolutionized image synthesis and manipulation, raising significant concerns about data misappropriation and intellectua…
SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
Quanjian Song, Donghao Zhou, Jingyu Lin +5
Recent text-to-image models have revolutionized image generation, but they still struggle with maintaining concept consistency across generated images. While existing works focus o…
MTS-Net: Dual-Enhanced Positional Multi-Head Self-Attention for 3D CT Diagnosis of May-Thurner Syndrome
Yixin Huang, Yiqi Jin, Ke Tao +6
May-Thurner Syndrome (MTS) is a vascular condition that affects over 20\% of the population and significantly increases the risk of iliofemoral deep venous thrombosis. Accurate and…