7 citations · 11 across the 7 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2022★ 3 cited
CPL: Counterfactual Prompt Learning for Vision and Language Models
Xuehai He, Diji Yang, Weixi Feng +7
Prompt tuning is a new few-shot transfer learning technique that only tunes the learnable prompt for pre-trained vision and language models such as CLIP. However, existing prompt t…
cs.CV2022★ 1 cited
Question Generation for Evaluating Cross-Dataset Shifts in Multi-modal Grounding
Arjun R. Akula
Visual question answering (VQA) is the multi-modal task of answering natural language questions about an input image. Through cross-dataset adaptation methods, it is possible to tr…
cs.CV2022
Discourse Analysis for Evaluating Coherence in Video Paragraph Captions
Arjun R Akula, Song-Chun Zhu
Video paragraph captioning is the task of automatically generating a coherent paragraph description of the actions in a video. Previous linguistic studies have demonstrated that co…