7 papers
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference
Xu Li, Yi Zheng, Mengyang Zhao +7
Large Vision-Language Models (LVLMs) typically require processing hundreds to thousands of visual tokens, leading to substantial inference overhead. Existing visual token pruning m…
Robust Embodied Perception in Dynamic Environments via Disentangled Weight Fusion
Juncen Guo, Xiaoguang Zhu, Jingyi Wu +4
Embodied perception systems face severe challenges of dynamic environment distribution drift when they continuously interact in open physical spaces. However, the existing domain i…
CalFuse: Multi-Modal Continual Learning via Feature Calibration and Parameter Fusion
Juncen Guo, Siao Liu, Xiaoguang Zhu +8
With the proliferation of multi-modal data in large-scale visual recognition systems, enabling models to continuously acquire knowledge from evolving data streams while preserving…
Privacy-Preserving Video Anomaly Detection: A Survey
Yang Liu, Siao Liu, Xiaoguang Zhu +7
Video Anomaly Detection (VAD) aims to automatically analyze spatiotemporal patterns in surveillance videos collected from open spaces to detect anomalous events that may cause harm…
Cross-channel Perception Learning for H&E-to-IHC Virtual Staining
Hao Yang, JianYu Wu, Run Fang +7
With the rapid development of digital pathology, virtual staining has become a key technology in multimedia medical information systems, offering new possibilities for the analysis…
A Survey on Video Analytics in Cloud-Edge-Terminal Collaborative Systems
Linxiao Gong, Hao Yang, Gaoyun Fang +7
The explosive growth of video data has driven the development of distributed video analytics in cloud-edge-terminal collaborative (CETC) systems, enabling efficient video processin…