11 papers · 1 filter
InstructVVT: Instruction-Driven Video Virtual Try-On without Auxiliary Spatial Priors
Dingbao Shao, Song Wu, Xinyu Chen +18
Video virtual try-on is a highly constrained editing task requiring the precise replacement of a target person's clothing while strictly preserving the original video's spatial str…
SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model
Zhennan Chen, Tianxing Shi, Pengcheng Xu +5
VAR has gained widespread popularity due to its next-scale prediction paradigm. However, it faces substantial performance bottlenecks when handling complex scenes with multiple obj…
Spiking Pyramid Wavelet Transformation for High-efficient and Low-energy Image Restoration
Chen Zhao, Xiantao Hu, Song Wu +5
Spiking neural networks (SNNs) have garnered significant interest in computer vision due to their potential for efficiency and biological inspiration. While spiking CNN-based metho…
TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On
Dingbao Shao, Song Wu, Shenyi Wang +9
Due to the scarcity of large-scale in-the-wild triplet data and the improper use of masks, the performance of video virtual try-on models remains limited. In this paper, we first i…
A training-free framework for high-fidelity appearance transfer via diffusion transformers
Shengrong Gu, Ye Wang, Song Wu +4
Diffusion Transformers (DiTs) excel at generation, but their global self-attention makes controllable, reference-image-based editing a distinct challenge. Unlike U-Nets, naively in…
FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction
Jiang Lin, Xinyu Chen, Song Wu +7
Controlling the spatial and semantic structure of diffusion-generated images remains a challenge. Existing methods like ControlNet rely on handcrafted condition maps and retraining…