2 citations · 2 across the 2 of their papers we have counts for
4 papers
Semantic-Aware Pretraining for Dense Video Captioning
Teng Wang, Zhu Liu, Feng Zheng +3
This report describes the details of our approach for the event dense-captioning task in ActivityNet Challenge 2021. We present a semantic-aware pretraining method for dense video…
Learning from Weakly-labeled Web Videos via Exploring Sub-Concepts
Kunpeng Li, Zizhao Zhang, Guanhang Wu +5
Learning visual knowledge from massive weakly-labeled web videos has attracted growing research interests thanks to the large corpus of easily accessible video data on the Internet…
DOPS: Learning to Detect 3D Objects and Predict their 3D Shapes
Mahyar Najibi, Guangda Lai, Abhijit Kundu +7
We propose DOPS, a fast single-stage 3D object detection method for LIDAR data. Previous methods often make domain-specific design decisions, for example projecting points into a b…
RetinaTrack: Online Single Stage Joint Detection and Tracking
Zhichao Lu, Vivek Rathod, Ronny Votel +1
Traditionally multi-object tracking and object detection are performed using separate systems with most prior works focusing exclusively on one of these aspects over the other. Tra…