6 papers
2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA
Zhiyu Wang, Xudong Kang, Shutao Li
Audio-based video object segmentation aims to locate and segment objects in videos conditioned on audio cues, requiring precise understanding of both appearance and motion. Recent…
Multimodal Prompt Alignment for Facial Expression Recognition
Fuyan Ma, Yiran He, Bin Sun +1
Prompt learning has been widely adopted to efficiently adapt vision-language models (VLMs) like CLIP for various downstream tasks. Despite their success, current VLM-based facial e…
NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment
Shuhao Han, Haotian Fan, Fangyuan Kong +112
This paper reports on the NTIRE 2025 challenge on Text to Image (T2I) generation model quality assessment, which will be held in conjunction with the New Trends in Image Restoratio…
DiffCL: A Diffusion-Based Contrastive Learning Framework with Semantic Alignment for Multimodal Recommendations
Qiya Song, Jiajun Hu, Lin Xiao +3
Multimodal recommendation systems integrate diverse multimodal information into the feature representations of both items and users, thereby enabling a more comprehensive modeling…
DrVideo: Document Retrieval Based Long Video Understanding
Ziyu Ma, Chenhui Gou, Hengcan Shi +4
Most of the existing methods for video understanding primarily focus on videos only lasting tens of seconds, with limited exploration of techniques for handling long videos. The in…
PVPUFormer: Probabilistic Visual Prompt Unified Transformer for Interactive Image Segmentation
Xu Zhang, Kailun Yang, Jiacheng Lin +3
Integration of diverse visual prompts like clicks, scribbles, and boxes in interactive image segmentation significantly facilitates users' interaction as well as improves interacti…