10 papers
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference
Xu Li, Yi Zheng, Mengyang Zhao +7
Large Vision-Language Models (LVLMs) typically require processing hundreds to thousands of visual tokens, leading to substantial inference overhead. Existing visual token pruning m…
PhysScene: A Scene Graph Dataset for Scientific Visual Reasoning in Physics Experiments
Minghao Zou, Qingtian Zeng, Shangkun Liu +5
Scene Graphs (SGs) provide structured representations of visual scenes by modeling objects and their pairwise relationships. Despite recent progress, existing datasets primarily fo…
Comparison Drives Preference: Reference-Aware Modeling for AI-Generated Video Quality Assessment
Minghao Zou, Gen Liu, Guanghui Yue +5
The rapid advancement of generative models has led to a growing volume of AI-generated videos, making the automatic quality assessment of such videos increasingly important. Existi…
Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2025
Jingzhe Ma, Meng Zhang, Jianlong Yu +34
Human identification at a distance (HID) is challenging because traditional biometric modalities such as face and fingerprints are often difficult to acquire in real-world scenario…
Depth-Guided Metric-Aware Temporal Consistency for Monocular Video Human Mesh Recovery
Jiaxin Cen, Xudong Mao, Guanghui Yue +4
Monocular video human mesh recovery faces fundamental challenges in maintaining metric consistency and temporal stability due to inherent depth ambiguities and scale uncertainties.…
Cross-Modal Scene Semantic Alignment for Image Complexity Assessment
Yuqing Luo, Yixiao Li, Jiang Liu +7
Image complexity assessment (ICA) is a challenging task in perceptual evaluation due to the subjective nature of human perception and the inherent semantic diversity in real-world…