From the 1 of 4 linked papers with an AI index.
4 papers
ME-IQA: Memory-Enhanced Image Quality Assessment via Re-Ranking
Kanglong Fan, Tianhe Wu, Wen Wen +6
The paper introduces ME‑IQA, a test‑time memory‑enhanced re‑ranking framework that leverages a memory bank of reasoning summaries to retrieve similar images, converts a vision‑lang…
TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning
Tao Wu, Li Yang, Gen Zhan +6
Enhancing the temporal understanding of Multimodal Large Language Models (MLLMs) is essential for advancing long-form video analysis, enabling tasks such as temporal localization,…
CaRDiff: Video Salient Object Ranking Chain of Thought Reasoning for Saliency Prediction with Diffusion
Yolo Yunlong Tang, Gen Zhan, Li Yang +2
Video saliency prediction aims to identify the regions in a video that attract human attention and gaze, driven by bottom-up features from the video and top-down processes like mem…
From Sight to Insight: Unleashing Eye-Tracking in Weakly Supervised Video Salient Object Detection
Qi Qin, Runmin Cong, Gen Zhan +2
The eye-tracking video saliency prediction (VSP) task and video salient object detection (VSOD) task both focus on the most attractive objects in video and show the result in the f…