2 papers
cs.CV2024
VG4D: Vision-Language Model Goes 4D Video Recognition
Zhichao Deng, Xiangtai Li, Xia Li +3
Understanding the real world through point cloud video is a crucial aspect of robotics and autonomous driving systems. However, prevailing methods for 4D point cloud recognition ha…
cs.CV2023
OV-VG: A Benchmark for Open-Vocabulary Visual Grounding
Chunlei Wang, Wenquan Feng, Xiangtai Li +5
Open-vocabulary learning has emerged as a cutting-edge research area, particularly in light of the widespread adoption of vision-based foundational models. Its primary objective is…