413 citations · 1.1k across the 43 of their papers we have counts for
4 papers · 1 filter
Chatting with Images for Introspective Visual Thinking
Junfei Wu, Jian Guan, Qiang Liu +4
Current large vision-language models (LVLMs) typically rely on text-only reasoning based on a single-pass visual encoding, which often leads to loss of fine-grained visual informat…
Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining
Han Huang, Yuqi Huo, Zijia Zhao +6
Multimodal large language models (MLLMs) have made significant strides by integrating visual and textual modalities. A critical factor in training MLLMs is the quality of image-tex…
Interpretable Multimodal Out-of-context Detection with Soft Logic Regularization
Huanhuan Ma, Jinghao Zhang, Qiang Liu +2
The rapid spread of information through mobile devices and media has led to the widespread of false or deceptive news, causing significant concerns in society. Among different type…
Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models
Junfei Wu, Qiang Liu, Ding Wang +4
Object hallucination has been an Achilles' heel which hinders the broader applications of large vision-language models (LVLMs). Object hallucination refers to the phenomenon that t…