From the 1 of 5 linked papers with an AI index.
5 papers
Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models
Zhiwei Yang, Yuanchen Wu, Nan Zhang +3
The paper proposes Scene Graph Thinking (SaGe), a method that equips multimodal large language models with explicit scene‑graph representations to enable fine‑grained, structured v…
Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models
Enguang Wang, Qiang Wang, Yuanchen Wu +5
While Multimodal Large Language Models (MLLMs) excel at vision-language tasks, the cost of their language-driven training on internal visual foundational competence remains unclear…
Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
Yuanchen Wu, Lu Zhang, Hang Yao +5
Large Vision-Language Models (LVLMs) have achieved impressive results across various cross-modal tasks. However, hallucinations, i.e., the models generating counterfactual response…
ToVE: Efficient Vision-Language Learning via Knowledge Transfer from Vision Experts
Yuanchen Wu, Junlong Du, Ke Yan +2
Vision-language (VL) learning requires extensive visual perception capabilities, such as fine-grained object recognition and spatial perception. Recent works typically rely on trai…
LaRE: Latent Reconstruction Error Based Method for Diffusion-Generated Image Detection
Yunpeng Luo, Junlong Du, Ke Yan +1
The evolution of Diffusion Models has dramatically improved image generation quality, making it increasingly difficult to differentiate between real and generated images. This deve…