2 papers
cs.CV2025
Can GPT tell us why these images are synthesized? Empowering Multimodal Large Language Models for Forensics
Yiran He, Yun Cao, Bowen Yang +1
The rapid development of generative AI facilitates content creation and makes image manipulation easier and more difficult to detect. While multimodal Large Language Models (LLMs)…
cs.CV2025
Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning
Yuqi Pang, Bowen Yang, Haoqin Tu +2
Although Large Language Models (LLMs) excel in reasoning and generation for language tasks, they are not specifically designed for multimodal challenges. Training Multimodal Large…