17 citations · 25 across the 4 of their papers we have counts for
4 papers
CREATE: A Benchmark for Chinese Short Video Retrieval and Title Generation
Ziqi Zhang, Yuxin Chen, Zongyang Ma +5
Previous works of video captioning aim to objectively describe the video's actual content, which lacks subjective and attractive expression, limiting its practical application scen…
Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge Distillation
Zongyang Ma, Guan Luo, Jin Gao +5
Open-vocabulary object detection aims to detect novel object categories beyond the training set. The advanced open-vocabulary two-stage detectors employ instance-level visual-to-vi…
Learning Target-aware Representation for Visual Tracking via Informative Interactions
Mingzhe Guo, Zhipeng Zhang, Heng Fan +4
We introduce a novel backbone architecture to improve target-perception ability of feature representation for tracking. Specifically, having observed that de facto frameworks perfo…
Differentiable Convolution Search for Point Cloud Processing
Xing Nie, Yongcheng Liu, Shaohong Chen +6
Exploiting convolutional neural networks for point cloud processing is quite challenging, due to the inherent irregular distribution and discrete shape representation of point clou…