85 citations · 102 across the 9 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2023★ 1 cited
Pro-Cap: Leveraging a Frozen Vision-Language Model for Hateful Meme Detection
Rui Cao, Ming Shan Hee, Adriel Kuek +3
Hateful meme detection is a challenging multimodal task that requires comprehension of both vision and language, as well as cross-modal interactions. Recent studies have tried to f…
cs.CV2023
Modularized Zero-shot VQA with Pre-trained Models
Rui Cao, Jing Jiang
Large-scale pre-trained models (PTMs) show great zero-shot capabilities. In this paper, we study how to leverage them for zero-shot visual question answering (VQA). Our approach is…