2 citations · 2 across the 4 of their papers we have counts for
4 papers
VG4D: Vision-Language Model Goes 4D Video Recognition
Zhichao Deng, Xiangtai Li, Xia Li +3
Understanding the real world through point cloud video is a crucial aspect of robotics and autonomous driving systems. However, prevailing methods for 4D point cloud recognition ha…
ModelNet-O: A Large-Scale Synthetic Dataset for Occlusion-Aware Point Cloud Classification
Zhongbin Fang, Xia Li, Xiangtai Li +2
Recently, 3D point cloud classification has made significant progress with the help of many datasets. However, these datasets do not reflect the incomplete nature of real-world poi…
Explore Human Parsing Modality for Action Recognition
Jinfu Liu, Runwei Ding, Yuhang Wen +4
Multimodal-based action recognition methods have achieved high success using pose and RGB modality. However, skeletons sequences lack appearance depiction and RGB images suffer irr…
GrowCLIP: Data-aware Automatic Model Growing for Large-scale Contrastive Language-Image Pre-training
Xinchi Deng, Han Shi, Runhui Huang +7
Cross-modal pre-training has shown impressive performance on a wide range of downstream tasks, benefiting from massive image-text pairs collected from the Internet. In practice, on…