computer vision

Towards Faithful Sentimental Image Captioning via Evidence-Aware Multi-Agent Reasoning

arXiv:2607.25789

summary

The paper introduces SEA-Cap, a multi‑agent system that extracts object‑level affective evidence from images and uses a generator, hallucination checker, and arbitrator to produce sentiment‑aware captions that are both emotionally appropriate and factually grounded.

Abstract

Sentimental Image Captioning (SIC) requires balancing emotional expression with visual fidelity. Existing methods often struggle with this trade-off, leading to hallucinations due to insufficient local grounding and the lack of sentimental verification mechanisms. To address these limitations, we propose SEA-Cap, a Sentiment-Evidence-Aware Multi-Agent System for faithful and evidence-grounded sentimental image captioning. SEA-Cap incorporates a Sentiment Evidence Miner that extracts structured, local affective cues to shift sentiment control from global attributes to verifiable object-level evidence. Leveraging this evidence, our framework orchestrates a collaborative workflow where a Generator, Hallucination Checker, and Arbitrator iteratively refine captions via a shared blackboard. By explicitly auditing generated content against mined visual evidence, SEA-Cap ensures both sentiment accuracy and factual consistency. Extensive experiments on two benchmark datasets demonstrate that SEA-Cap effectively mitigates hallucinations and achieves state-of-the-art performance.

Topics & keywords

#sentimental image captioning#evidence-aware reasoning#multi-agent systems#hallucination mitigation#visual groundingsentiment evidence minerhallucination checkerarbitratorblackboard architectureaffective cuesimage captioning