6 citations · 40 across the 46 of their papers we have counts for
Showing 2024 · cs.CVShow all
3 papers · 2 filters
cs.CV2024
Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining
Han Huang, Yuqi Huo, Zijia Zhao +6
Multimodal large language models (MLLMs) have made significant strides by integrating visual and textual modalities. A critical factor in training MLLMs is the quality of image-tex…
cs.CV2024
Interpretable Multimodal Out-of-context Detection with Soft Logic Regularization
Huanhuan Ma, Jinghao Zhang, Qiang Liu +2
The rapid spread of information through mobile devices and media has led to the widespread of false or deceptive news, causing significant concerns in society. Among different type…
cs.CV2024
Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models
Junfei Wu, Qiang Liu, Ding Wang +4
Object hallucination has been an Achilles' heel which hinders the broader applications of large vision-language models (LVLMs). Object hallucination refers to the phenomenon that t…