12 papers
MER-Bench: A Comprehensive Benchmark for Multimodal Meme Reappraisal
Yiqi Nie, Fei Wang, Junjie Chen +5
Memes represent a tightly coupled, multimodal form of social expression, in which visual context and overlaid text jointly convey nuanced affect and commentary. Inspired by cogniti…
Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization
Cailing Han, Zhangbin Li, Jinxing Zhou +5
Point-level weakly-supervised temporal sentiment localization (P-WTSL) aims to detect sentiment-relevant segments in untrimmed multimodal videos using timestamp sentiment annotatio…
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
Zhenxing Zhang, Yaxiong Wang, Lechao Cheng +3
We present ASAP, a new framework for detecting and grounding multi-modal media manipulation (DGM4).Upon thorough examination, we observe that accurate fine-grained cross-modal sema…
EmoSEM: Segment and Explain Emotion Stimuli in Visual Art
Jing Zhang, Dan Guo, Zhangbin Li +1
This paper focuses on a key challenge in visual emotion understanding: given an art image, the model pinpoints pixel regions that trigger a specific human emotion, and generates li…
SSAM: Self-Supervised Association Modeling for Test-Time Adaption
Yaxiong Wang, Zhenqiang Zhang, Lechao Cheng +3
Test-time adaption (TTA) has witnessed important progress in recent years, the prevailing methods typically first encode the image and the text and design strategies to model the a…
Scene-Text Grounding for Text-Based Video Question Answering
Sheng Zhou, Junbin Xiao, Xun Yang +5
Existing efforts in text-based video question answering (TextVideoQA) are criticized for their opaque decisionmaking and heavy reliance on scene-text recognition. In this paper, we…