Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Knowledge is Power: Advancing Few-shot Action Recognition with Multimodal Semantics from MLLMs
Jiazheng Xing, Chao Xu, Hangjie Yuan +4
Multimodal Large Language Models (MLLMs) have propelled the field of few-shot action recognition (FSAR). However, preliminary explorations in this area primarily focus on generatin…
cs.CV2024
MA-FSAR: Multimodal Adaptation of CLIP for Few-Shot Action Recognition
Jiazheng Xing, Chao Xu, Mengmeng Wang +5
Applying large-scale vision-language pre-trained models like CLIP to few-shot action recognition (FSAR) can significantly enhance both performance and efficiency. While several stu…
cs.CV2024
Camera-based 3D Semantic Scene Completion with Sparse Guidance Network
Jianbiao Mei, Yu Yang, Mengmeng Wang +5
Semantic scene completion (SSC) aims to predict the semantic occupancy of each voxel in the entire 3D scene from limited observations, which is an emerging and critical task for au…