2 papers
cs.CV2026
Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning
Kazi Sajeed Mehrab, Hani Alomari, Najibul Haque Sarker +4
Multimodal large language models (MLLMs) ground whole objects well from free-form language queries, but they struggle when the query names a part rather than the object. We trace t…
cs.AI2024
MetaSumPerceiver: Multimodal Multi-Document Evidence Summarization for Fact-Checking
Ting-Chih Chen, Chia-Wei Tang, Chris Thomas
Fact-checking real-world claims often requires reviewing multiple multimodal documents to assess a claim's truthfulness, which is a highly laborious and time-consuming task. In thi…