3 papers
cs.CV2025
RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence
Xuming He, Zehao Fan, Hengjia Li +7
Recent advances in video generation have enabled the synthesis of videos with strong temporal consistency and impressive visual quality, marking a crucial step toward vision founda…
cs.CV2025
Video Virtual Try-on with Conditional Diffusion Transformer Inpainter
Cheng Zou, Senlin Cheng, Bolei Xu +4
Video virtual try-on aims to naturally fit a garment to a target person in consecutive video frames. It is a challenging task, on the one hand, the output video should be in good s…
cs.CV2024
SPT: Sequence Prompt Transformer for Interactive Image Segmentation
Senlin Cheng, Haopeng Sun
Interactive segmentation aims to extract objects of interest from an image based on user-provided clicks. In real-world applications, there is often a need to segment a series of i…