9 papers
Preference-Aware Rubric Learning for Personalized Evaluation
Yilun Qiu, Xiaoyan Zhao, Yang Zhang +7
As Large Language Models (LLMs) evolve from general-purpose assistants to user-centric agents, personalization has become central to aligning model behavior with individual prefere…
AgentCVR: Active Multi-Agent Cross-Video Reasoning via Script-Simulated Reinforcement Learning
Yilun Qiu, Jiahe Wang, Cilin Yan +4
Cross-Video Reasoning (CVR) has emerged as a critical frontier in multimodal intelligence, requiring models to retrieve, align, and aggregate evidence distributed across multiple v…
Rethinking the Global Knowledge of CLIP in Training-Free Open-Vocabulary Semantic Segmentation
Jingyun Wang, Cilin Yan, Guoliang Kang
Recent works modify CLIP to perform open-vocabulary semantic segmentation in a training-free manner (TF-OVSS). In vanilla CLIP, patch-wise image representations mainly encode homog…
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
Bob Zhang, Haoran Li, Tao Zhang +5
Multimodal Large Language Models (MLLMs) perform well in single-image visual grounding but struggle with real-world tasks that demand cross-image reasoning and multi-modal instruct…
CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
Jingyao Li, Jingyun Wang, Molin Tan +6
Cross-Video Reasoning (CVR) presents a significant challenge in video understanding, which requires simultaneous understanding of multiple videos to aggregate and compare informati…
LTCA: Long-range Temporal Context Attention for Referring Video Object Segmentation
Cilin Yan, Jingyun Wang, Guoliang Kang
Referring Video Segmentation (RVOS) aims to segment objects in videos given linguistic expressions. The key to solving RVOS is to extract long-range temporal context information fr…