24 citations · 66 across the 26 of their papers we have counts for
Showing 2023 · cs.CVShow all
2 papers · 2 filters
cs.CV2023★ 1 cited
ZEETAD: Adapting Pretrained Vision-Language Model for Zero-Shot End-to-End Temporal Action Detection
Thinh Phan, Khoa Vo, Duy Le +3
Temporal action detection (TAD) involves the localization and classification of action instances within untrimmed videos. While standard TAD follows fully supervised learning with…
cs.CV2023★ 1 cited
Open-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene Representation
Kashu Yamazaki, Taisei Hanyu, Khoa Vo +5
Precise 3D environmental mapping is pivotal in robotics. Existing methods often rely on predefined concepts during training or are time-intensive when generating semantic maps. Thi…