2 papers
cs.CV2026
InstructVVT: Instruction-Driven Video Virtual Try-On without Auxiliary Spatial Priors
Dingbao Shao, Song Wu, Xinyu Chen +18
Video virtual try-on is a highly constrained editing task requiring the precise replacement of a target person's clothing while strictly preserving the original video's spatial str…
cs.CV2026
A training-free framework for high-fidelity appearance transfer via diffusion transformers
Shengrong Gu, Ye Wang, Song Wu +4
Diffusion Transformers (DiTs) excel at generation, but their global self-attention makes controllable, reference-image-based editing a distinct challenge. Unlike U-Nets, naively in…