2 papers
cs.CV2026
InstructVVT: Instruction-Driven Video Virtual Try-On without Auxiliary Spatial Priors
Dingbao Shao, Song Wu, Xinyu Chen +18
Video virtual try-on is a highly constrained editing task requiring the precise replacement of a target person's clothing while strictly preserving the original video's spatial str…
cs.CV2026
EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory
Weitao Chen, Hu Jiaxin, Xie Tianyidan +15
Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. Howev…