Search papers, labs, and topics across Lattice.
This paper introduces SEA-Cap, a novel Sentiment-Evidence-Aware Multi-Agent System designed to enhance Sentimental Image Captioning (SIC) by integrating local affective cues and a collaborative workflow among agents. By shifting sentiment control from global attributes to verifiable object-level evidence, SEA-Cap effectively reduces hallucinations and improves the accuracy of emotional expression in generated captions. Extensive evaluations on benchmark datasets reveal that SEA-Cap not only mitigates hallucinations but also achieves state-of-the-art performance in SIC tasks.
Hallucinations in sentimental image captions can be drastically reduced by grounding sentiment in verifiable object-level evidence.
Sentimental Image Captioning (SIC) requires balancing emotional expression with visual fidelity. Existing methods often struggle with this trade-off, leading to hallucinations due to insufficient local grounding and the lack of sentimental verification mechanisms. To address these limitations, we propose SEA-Cap, a Sentiment-Evidence-Aware Multi-Agent System for faithful and evidence-grounded sentimental image captioning. SEA-Cap incorporates a Sentiment Evidence Miner that extracts structured, local affective cues to shift sentiment control from global attributes to verifiable object-level evidence. Leveraging this evidence, our framework orchestrates a collaborative workflow where a Generator, Hallucination Checker, and Arbitrator iteratively refine captions via a shared blackboard. By explicitly auditing generated content against mined visual evidence, SEA-Cap ensures both sentiment accuracy and factual consistency. Extensive experiments on two benchmark datasets demonstrate that SEA-Cap effectively mitigates hallucinations and achieves state-of-the-art performance.