4 papers
Knowledge is Power: Advancing Few-shot Action Recognition with Multimodal Semantics from MLLMs
Jiazheng Xing, Chao Xu, Hangjie Yuan +4
Multimodal Large Language Models (MLLMs) have propelled the field of few-shot action recognition (FSAR). However, preliminary explorations in this area primarily focus on generatin…
CFSum: A Transformer-Based Multi-Modal Video Summarization Framework With Coarse-Fine Fusion
Yaowei Guo, Jiazheng Xing, Xiaojun Hou +5
Video summarization, by selecting the most informative and/or user-relevant parts of original videos to create concise summary videos, has high research value and consumer demand i…
Visual Object Tracking across Diverse Data Modalities: A Review
Mengmeng Wang, Teli Ma, Shuo Xin +5
Visual Object Tracking (VOT) is an attractive and significant research area in computer vision, which aims to recognize and track specific targets in video sequences where the targ…
MA-FSAR: Multimodal Adaptation of CLIP for Few-Shot Action Recognition
Jiazheng Xing, Chao Xu, Mengmeng Wang +5
Applying large-scale vision-language pre-trained models like CLIP to few-shot action recognition (FSAR) can significantly enhance both performance and efficiency. While several stu…