28 papers
Aggregating Visual Information with Optimal Transport for VideoLM Token Compression
Wenti Yin, Xiaotian Han, Junyuan Shang +5
Video language models process videos as dense visual-token sequences with substantial representational redundancy. Compressing these sequences is therefore essential for reducing t…
Beyond Hazard Resemblance: Contrastive Event Adjudication for Training-Free Video Anomaly Detection
Wenti Yin, Xiang Wang, Huaxin Zhang +4
Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries from target-domain annotation…
Clarity Contrast and Similarity Selection for Multi-Focus Image Fusion
Yicheng Zhang, Haoyou Deng, Zhiqiang Li +3
Multi-focus image fusion (MFIF) aims to generate an all-in-focus image from multiple images of the same scene focused at different regions. Most existing deep learning-based method…
Achieving Text-based Person Retrieval with Any Granularity
Jialong Zuo, Hanyu Zhou, Dongyue Wu +5
Text-based person retrieval faces a critical but under-explored challenge: the inherent uncertainty of query granularity in real-world scenarios. This paper introduces a new paradi…
FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling
Jialong Zuo, Haotong Zuo, Shiwei Zhang +5
Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-form, multi-scene visual nar…
Selecting Samples on Graphs: A Unified Dataset Pruning Framework for Lossless Training Acceleration
Dongyue Wu, Zilin Guo, Xiaoyu Li +4
The rapid growth of modern training datasets has significantly increased computational cost, motivating dataset pruning~(DP) methods which retain only a subset of informative sampl…