5 papers
From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer Learning
Yang Liu, Qianqian Xu, Peisong Wen +3
Recent studies have made notable progress in video representation learning by transferring image-pretrained models to video tasks, typically with complex temporal modules and video…
Guiding Diffusion-based Reconstruction with Contrastive Signals for Balanced Visual Representation
Boyu Han, Qianqian Xu, Shilong Bao +4
The limited understanding capacity of the visual encoder in Contrastive Language-Image Pre-training (CLIP) has become a key bottleneck for downstream performance. This capacity inc…
BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response Deviation
Feiran Li, Qianqian Xu, Shilong Bao +4
This paper investigates the challenging task of detecting backdoored text-to-image models under black-box settings and introduces a novel detection framework BlackMirror. Existing…
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
Yang Liu, Xilin Zhao, Peisong Wen +2
Recent progress in video generation has led to impressive visual quality, yet current models still struggle to produce results that align with real-world physical principles. To th…
Mano Technical Report
Tianyu Fu, Anyang Su, Chenxu Zhao +20
Graphical user interfaces (GUIs) are the primary medium for human-computer interaction, yet automating GUI interactions remains challenging due to the complexity of visual elements…