Showing 2025Show all
3 papers · 1 filter
cs.CV2025
Language-Guided Graph Representation Learning for Video Summarization
Wenrui Li, Wei Han, Hengyu Man +3
With the rapid growth of video content on social media, video summarization has become a crucial task in multimedia processing. However, existing methods face challenges in capturi…
cs.CV2025
Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
Yuyang Liu, Qiuhe Hong, Linlan Huang +6
Vision-language models (VLMs), spanning predictive architectures to generative Multimodal Large Language Models (MLLMs), have revolutionized artificial intelligence through powerfu…
cs.CV2025
Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning
Wenrui Li, Penghong Wang, Xingtao Wang +3
Audio-visual zero-shot learning (ZSL) has been extensively researched for its capability to classify video data from unseen classes during training. Nevertheless, current methodolo…