collaborators

10 papers

cs.CV2026

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference

Xu Li, Yi Zheng, Mengyang Zhao +7

Large Vision-Language Models (LVLMs) typically require processing hundreds to thousands of visual tokens, leading to substantial inference overhead. Existing visual token pruning m…

cs.CV2026

PhysScene: A Scene Graph Dataset for Scientific Visual Reasoning in Physics Experiments

Minghao Zou, Qingtian Zeng, Shangkun Liu +5

Scene Graphs (SGs) provide structured representations of visual scenes by modeling objects and their pairwise relationships. Despite recent progress, existing datasets primarily fo…

cs.CV2026

Comparison Drives Preference: Reference-Aware Modeling for AI-Generated Video Quality Assessment

Minghao Zou, Gen Liu, Guanghui Yue +5

The rapid advancement of generative models has led to a growing volume of AI-generated videos, making the automatic quality assessment of such videos increasingly important. Existi…

cs.CV2026

Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2025

Jingzhe Ma, Meng Zhang, Jianlong Yu +34

Human identification at a distance (HID) is challenging because traditional biometric modalities such as face and fingerprints are often difficult to acquire in real-world scenario…

cs.CV2026

Depth-Guided Metric-Aware Temporal Consistency for Monocular Video Human Mesh Recovery

Jiaxin Cen, Xudong Mao, Guanghui Yue +4

Monocular video human mesh recovery faces fundamental challenges in maintaining metric consistency and temporal stability due to inherent depth ambiguities and scale uncertainties.…

cs.CV2025

Cross-Modal Scene Semantic Alignment for Image Complexity Assessment

Yuqing Luo, Yixiao Li, Jiang Liu +7

Image complexity assessment (ICA) is a challenging task in perceptual evaluation due to the subjective nature of human perception and the inherent semantic diversity in real-world…