3 papers
cs.CV2026
GRACE: Boosting Video MLLMs with Grounded Action-Centric Evidence for Viewer Sentiment Prediction
Ruoxuan Yang, Tieyuan Chen, Xiaofeng Huang +6
Viewer sentiment prediction in video advertisements aims to infer the latent affective response evoked in the audience. To bridge the gap between what is shown and what is felt, mo…
cs.CL2025
From Clicks to Preference: A Multi-stage Alignment Framework for Generative Query Suggestion in Conversational System
Junhao Yin, Haolin Wang, Peng Bao +2
Generative query suggestion using large language models offers a powerful way to enhance conversational systems, but aligning outputs with nuanced user preferences remains a critic…
cs.CV2025
ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models
Jingwei Yi, Junhao Yin, Ju Xu +4
Vision-Language Models (VLMs) have demonstrated remarkable capabilities in understanding multimodal inputs and have been widely integrated into Retrieval-Augmented Generation (RAG)…