2 papers
cs.CV2026
Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models
Yuriel Ryan, Hei Man Ip, Adriel Kuek +2
Current vision language models face hallucination and robustness issues against ambiguous or corrupted modalities. We hypothesize that these issues can be addressed by exploiting t…
cs.CV2025
A Closer Look at Dynamic Scene Graph Generation In the Era of Multimodal Large Language Models
Xuanming Cui, Jaiminkumar Ashokbhai Bhoi, Chionh Wei Peng +2
Dynamic Scene Graph Generation (DSGG) aims to capture objects and their evolving relations in videos. Despite recent progress, the practicality and quality of generated scene graph…