7 papers
Aggregating Visual Information with Optimal Transport for VideoLM Token Compression
Wenti Yin, Xiaotian Han, Junyuan Shang +5
Video language models process videos as dense visual-token sequences with substantial representational redundancy. Compressing these sequences is therefore essential for reducing t…
Beyond Hazard Resemblance: Contrastive Event Adjudication for Training-Free Video Anomaly Detection
Wenti Yin, Xiang Wang, Huaxin Zhang +4
Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries from target-domain annotation…
Clarity Contrast and Similarity Selection for Multi-Focus Image Fusion
Yicheng Zhang, Haoyou Deng, Zhiqiang Li +3
Multi-focus image fusion (MFIF) aims to generate an all-in-focus image from multiple images of the same scene focused at different regions. Most existing deep learning-based method…
Diffusion Models are Open-World Affordance Learners: Leveraging Generative Priors for 3D Affordance Learning
Hanqing Wang, Zhenhao Zhang, Kaiyang Ji +12
3D affordance grounding aims to understand how diverse objects can be manipulated, making it a cornerstone of embodied interaction. However, prior works struggle to generalize to o…
VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model
Hanqing Wang, Mingyu Liu, Xiaoyu Chen +9
3D affordance grounding aims to highlight the actionable regions on 3D objects, which is crucial for robotic manipulation. Previous research primarily focused on learning affordanc…
Learning to Tell Apart: Weakly Supervised Video Anomaly Detection via Disentangled Semantic Alignment
Wenti Yin, Huaxin Zhang, Xiang Wang +7
Recent advancements in weakly-supervised video anomaly detection have achieved remarkable performance by applying the multiple instance learning paradigm based on multimodal founda…