19 papers
InstructVVT: Instruction-Driven Video Virtual Try-On without Auxiliary Spatial Priors
Dingbao Shao, Song Wu, Xinyu Chen +18
Video virtual try-on is a highly constrained editing task requiring the precise replacement of a target person's clothing while strictly preserving the original video's spatial str…
Magnetism in antiperovskite (Li\textit{M})\textit{Ch}O (\textit{M} = Fe, Mn, Co; \textit{Ch} = S, Se) diluted magnets with fixed 1/3 filling: the key role of magnetic anisotropy
Jieyuan Zheng, Frederik L. Carstens, Lennart Singer +7
We report the magnetic properties of a series of lithium-rich antiperovskites (Li)O ( = Fe, Co, Mn and = Se, S) where transition metal and lithium ions are randoml…
SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model
Zhennan Chen, Tianxing Shi, Pengcheng Xu +5
VAR has gained widespread popularity due to its next-scale prediction paradigm. However, it faces substantial performance bottlenecks when handling complex scenes with multiple obj…
SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion
Xinyu Chen, Yuyi Qian, Jiang Lin +9
Video object insertion requires ensuring spatio-temporal coherence and interactive realism, extending far beyond simple content placement. However, current approaches are often hin…
Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance
Song Wu, Xinyu Chen, Qian Wang +3
Video editing poses a significant challenge. While a series of tuning-free methods circumvent the need for extensive data collection and model training, they often underutilize the…
TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On
Dingbao Shao, Song Wu, Shenyi Wang +9
Due to the scarcity of large-scale in-the-wild triplet data and the improper use of masks, the performance of video virtual try-on models remains limited. In this paper, we first i…