4 citations · 6 across the 4 of their papers we have counts for
4 papers
Mug-STAN: Adapting Image-Language Pretrained Models for General Video Understanding
Ruyang Liu, Jingjia Huang, Wei Gao +2
Large-scale image-language pretrained models, e.g., CLIP, have demonstrated remarkable proficiency in acquiring general multi-modal knowledge through web-scale image-text data. Des…
Efficient Test-Time Adaptation for Super-Resolution with Second-Order Degradation and Reconstruction
Zeshuai Deng, Zhuokun Chen, Shuaicheng Niu +3
Image super-resolution (SR) aims to learn a mapping from low-resolution (LR) to high-resolution (HR) using paired HR-LR training images. Conventional SR methods typically gather th…
Hard Sample Matters a Lot in Zero-Shot Quantization
Huantong Li, Xiangmiao Wu, Fanbing Lv +5
Zero-shot quantization (ZSQ) is promising for compressing and accelerating deep neural networks when the data for training full-precision models are inaccessible. In ZSQ, network q…
Revisiting Temporal Modeling for CLIP-based Image-to-Video Knowledge Transferring
Ruyang Liu, Jingjia Huang, Ge Li +3
Image-text pretrained models, e.g., CLIP, have shown impressive general multi-modal knowledge learned from large-scale image-text data pairs, thus attracting increasing attention f…