3 papers
cs.CV2026
Towards Faithful Sentimental Image Captioning via Evidence-Aware Multi-Agent Reasoning
Tiecheng Cai, Zexian Yang, Chao Chen +2
The paper introduces SEA-Cap, a multi‑agent system that extracts object‑level affective evidence from images and uses a generator, hallucination checker, and arbitrator to produce…
cs.CV2025
Uneven Event Modeling for Partially Relevant Video Retrieval
Sa Zhu, Huashan Chen, Wanqian Zhang +4
Given a text query, partially relevant video retrieval (PRVR) aims to retrieve untrimmed videos containing relevant moments, wherein event modeling is crucial for partitioning the…
cs.CV2025
Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning
Zexian Yang, Dian Li, Dayan Wu +2
Despite significant advancements in multimodal reasoning tasks, existing Large Vision-Language Models (LVLMs) are prone to producing visually ungrounded responses when interpreting…