Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
GRACE: Boosting Video MLLMs with Grounded Action-Centric Evidence for Viewer Sentiment Prediction
Ruoxuan Yang, Tieyuan Chen, Xiaofeng Huang +6
Viewer sentiment prediction in video advertisements aims to infer the latent affective response evoked in the audience. To bridge the gap between what is shown and what is felt, mo…
cs.CV2025
ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models
Jingwei Yi, Junhao Yin, Ju Xu +4
Vision-Language Models (VLMs) have demonstrated remarkable capabilities in understanding multimodal inputs and have been widely integrated into Retrieval-Augmented Generation (RAG)…