1 citations · 1 across the 2 of their papers we have counts for
3 papers · 1 filter
Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models
Yuriel Ryan, Hei Man Ip, Adriel Kuek +2
Current vision language models face hallucination and robustness issues against ambiguous or corrupted modalities. We hypothesize that these issues can be addressed by exploiting t…
A Closer Look at Dynamic Scene Graph Generation In the Era of Multimodal Large Language Models
Xuanming Cui, Jaiminkumar Ashokbhai Bhoi, Chionh Wei Peng +2
Dynamic Scene Graph Generation (DSGG) aims to capture objects and their evolving relations in videos. Despite recent progress, the practicality and quality of generated scene graph…
Pro-Cap: Leveraging a Frozen Vision-Language Model for Hateful Meme Detection
Rui Cao, Ming Shan Hee, Adriel Kuek +3
Hateful meme detection is a challenging multimodal task that requires comprehension of both vision and language, as well as cross-modal interactions. Recent studies have tried to f…