1 citations · 1 across the 5 of their papers we have counts for
8 papers
Embodied Multimedia: A Tutorial
Yang Liu, Wei Zuo, Guanwei Zhao +4
Traditional multimedia technology has been built around optimizing content delivery for human observers, from perceptually driven compression standards to human-centric quality met…
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference
Xu Li, Yi Zheng, Mengyang Zhao +7
Large Vision-Language Models (LVLMs) typically require processing hundreds to thousands of visual tokens, leading to substantial inference overhead. Existing visual token pruning m…
Robust Embodied Perception in Dynamic Environments via Disentangled Weight Fusion
Juncen Guo, Xiaoguang Zhu, Jingyi Wu +4
Embodied perception systems face severe challenges of dynamic environment distribution drift when they continuously interact in open physical spaces. However, the existing domain i…
Cross-channel Perception Learning for H&E-to-IHC Virtual Staining
Hao Yang, JianYu Wu, Run Fang +7
With the rapid development of digital pathology, virtual staining has become a key technology in multimedia medical information systems, offering new possibilities for the analysis…
Adaptive Weighted Parameter Fusion with CLIP for Class-Incremental Learning
Juncen Guo, Xiaoguang Zhu, Liangyu Teng +4
Class-incremental Learning (CIL) enables the model to incrementally absorb knowledge from new classes and build a generic classifier across all previously encountered classes. When…
CalFuse: Multi-Modal Continual Learning via Feature Calibration and Parameter Fusion
Juncen Guo, Siao Liu, Xiaoguang Zhu +8
With the proliferation of multi-modal data in large-scale visual recognition systems, enabling models to continuously acquire knowledge from evolving data streams while preserving…