5 papers
Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
Yuyang Liu, Qiuhe Hong, Linlan Huang +6
Vision-language models (VLMs), spanning predictive architectures to generative Multimodal Large Language Models (MLLMs), have revolutionized artificial intelligence through powerfu…
All-in-One Video Restoration under Smoothly Evolving Unknown Weather Degradations
Wenrui Li, Hongtao Chen, Yao Xiao +4
All-in-one image restoration aims to recover clean images from diverse unknown degradations using a single model. But extending this task to videos faces unique challenges. Existin…
Language-Guided Graph Representation Learning for Video Summarization
Wenrui Li, Wei Han, Hengyu Man +3
With the rapid growth of video content on social media, video summarization has become a crucial task in multimedia processing. However, existing methods face challenges in capturi…
Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning
Wenrui Li, Penghong Wang, Xingtao Wang +3
Audio-visual zero-shot learning (ZSL) has been extensively researched for its capability to classify video data from unseen classes during training. Nevertheless, current methodolo…
Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body
Zeqing Wang, Qingyang Ma, Wentao Wan +3
Recent improvements in visual synthesis have significantly enhanced the depiction of generated human photos, which are pivotal due to their wide applicability and demand. Nonethele…