4 citations · 4 across the 1 of their papers we have counts for
2 papers
cs.CV2026★ 4 cited
ForgeryGPT: A Multimodal LLM for Interpretable Image Forgery Detection and Localization
Fanrui Zhang, Jiawei Liu, Jiaying Zhu +4
Multimodal Large Language Models (MLLMs), such as GPT4o, have shown strong capabilities in visual reasoning and explanation generation. However, despite these strengths, they face…
cs.CV2025
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
Fan Lu, Wei Wu, Kecheng Zheng +7
Generating detailed captions comprehending text-rich visual content in images has received growing attention for Large Vision-Language Models (LVLMs). However, few studies have dev…