13 citations · 23 across the 5 of their papers we have counts for
5 papers · 1 filter
MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning
Bang Yang, Fenglin Liu, Xian Wu +3
Supervised visual captioning models typically require a large scale of images or videos paired with descriptions in a specific language (i.e., the vision-caption pairs) for trainin…
Strip-MLP: Efficient Token Interaction for Vision MLP
Guiping Cao, Shengda Luo, Wenjian Huang +4
Token interaction operation is one of the core modules in MLP-based models to exchange and aggregate information between different spatial locations. However, the power of token in…
Towards Efficient Task-Driven Model Reprogramming with Foundation Models
Shoukai Xu, Jiangchao Yao, Ran Luo +5
Vision foundation models exhibit impressive power, benefiting from the extremely large model capacity and broad training data. However, in practice, downstream scenarios may only s…
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding
Linhui Xiao, Xiaoshan Yang, Fang Peng +3
Visual Grounding (VG) is a crucial topic in the field of vision and language, which involves locating a specific region described by expressions within an image. To reduce the reli…
Boost Test-Time Performance with Closed-Loop Inference
Shuaicheng Niu, Jiaxiang Wu, Yifan Zhang +6
Conventional deep models predict a test sample with a single forward propagation, which, however, may not be sufficient for predicting hard-classified samples. On the contrary, we…