activity
20242026
collaborators

30 papers

cs.CV2026

Aggregating Visual Information with Optimal Transport for VideoLM Token Compression

Wenti Yin, Xiaotian Han, Junyuan Shang +5

Video language models process videos as dense visual-token sequences with substantial representational redundancy. Compressing these sequences is therefore essential for reducing t…

cs.CV2026

Beyond Hazard Resemblance: Contrastive Event Adjudication for Training-Free Video Anomaly Detection

Wenti Yin, Xiang Wang, Huaxin Zhang +4

Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries from target-domain annotation…

cs.CV2026

Clarity Contrast and Similarity Selection for Multi-Focus Image Fusion

Yicheng Zhang, Haoyou Deng, Zhiqiang Li +3

Multi-focus image fusion (MFIF) aims to generate an all-in-focus image from multiple images of the same scene focused at different regions. Most existing deep learning-based method…

cs.CV2026

Achieving Text-based Person Retrieval with Any Granularity

Jialong Zuo, Hanyu Zhou, Dongyue Wu +5

Text-based person retrieval faces a critical but under-explored challenge: the inherent uncertainty of query granularity in real-world scenarios. This paper introduces a new paradi…

cs.CV2026

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling

Jialong Zuo, Haotong Zuo, Shiwei Zhang +5

Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-form, multi-scene visual nar…

cs.LG2026

Selecting Samples on Graphs: A Unified Dataset Pruning Framework for Lossless Training Acceleration

Dongyue Wu, Zilin Guo, Xiaoyu Li +4

The rapid growth of modern training datasets has significantly increased computational cost, motivating dataset pruning~(DP) methods which retain only a subset of informative sampl…