17 papers
PRISMR: Overcoming Parse Collapse in Multimodal Listwise Ranking via Parameterized Representation Internalization
Hao Jiang, Xin Li, Annan Wang +4
Generative listwise ranking with Large Multimodal Models (LMMs) aims to capture global list context in a single forward pass, but its effectiveness degrades in long-context multimo…
ConceptSeg-R1: Segment Any Concept via Meta-Reinforcement Learning
Yuan Zhao, Youwei Pang, Jiaming Zuo +10
Recent progress in promptable segmentation has shifted visual perception from object-level localization toward concept-level understanding. However, the notion of a concept remains…
R4-CGQA: Retrieval-based Vision Language Models for Computer Graphics Image Quality Assessment
Zhuangzi Li, Jian Jin, Shilv Cai +1
Immersive Computer Graphics (CGs) rendering has become ubiquitous in modern daily life. However, comprehensively evaluating CG quality remains challenging for two reasons: First, e…
TPIFM: A Task-Aware Model for Evaluating Perceptual Interaction Fluency in Remote AR Collaboration
Jiarun Song, Ninghao Wan, Fuzheng Yang +1
Remote Collaborative Augmented Reality (RCAR) enables geographically distributed users to collaborate by integrating virtual and physical environments. However, because RCAR relies…
From Perception to Cognition: How Latency Affects Interaction Fluency and Social Presence in VR Conferencing
Jiarun Song, Ninghao Wan, FuZheng Yang +1
Virtual reality (VR) conferencing has the potential to provide geographically dispersed users with an immersive environment, enabling rich social interactions and user experience u…
Enhancing Blind Video Quality Assessment with Rich Quality-aware Features
Wei Sun, Linhan Cao, Jun Jia +4
Blind video quality assessment (BVQA) is a highly challenging task due to the intrinsic complexity of video content and visual distortions, especially given the high popularity of…