32 citations · 120 across the 21 of their papers we have counts for
27 papers · 1 filter
Towards General Visual-Linguistic Face Forgery Detection(V2)
Ke Sun, Shen Chen, Taiping Yao +5
Face manipulation techniques have achieved significant advances, presenting serious challenges to security and social trust. Recent works demonstrate that leveraging multimodal mod…
Parameter and Computation Efficient Transfer Learning for Vision-Language Pre-trained Models
Qiong Wu, Wei Yu, Yiyi Zhou +3
With ever increasing parameters and computation, vision-language pre-trained (VLP) models exhibit prohibitive expenditure in downstream task adaption. Recent endeavors mainly focus…
AutoDiffusion: Training-Free Optimization of Time Steps and Architectures for Automated Diffusion Model Acceleration
Lijiang Li, Huixia Li, Xiawu Zheng +7
Diffusion models are emerging expressive generative models, in which a large number of time steps (inference steps) are required for a single image generation. To accelerate such t…
Towards Unified Token Learning for Vision-Language Tracking
Yaozong Zheng, Bineng Zhong, Qihua Liang +3
In this paper, we present a simple, flexible and effective vision-language (VL) tracking pipeline, termed \textbf{MMTrack}, which casts VL tracking as a token generation task. Trad…
Improving Human-Object Interaction Detection via Virtual Image Learning
Shuman Fang, Shuai Liu, Jie Li +3
Human-Object Interaction (HOI) detection aims to understand the interactions between humans and objects, which plays a curtail role in high-level semantic understanding tasks. Howe…
Beyond First Impressions: Integrating Joint Multi-modal Cues for Comprehensive 3D Representation
Haowei Wang, Jiji Tang, Jiayi Ji +8
In recent years, 3D understanding has turned to 2D vision-language pre-trained models to overcome data scarcity challenges. However, existing methods simply transfer 2D alignment s…