9 citations · 9 across the 4 of their papers we have counts for
5 papers
Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction
Mingda Jia, Weiliang Meng, Zenghuang Fu +7
Dense video captioning jointly localizes and captions salient events in untrimmed videos. Recent methods primarily focus on leveraging additional prior knowledge and advanced multi…
Image Recognition with Online Lightweight Vision Transformer: A Survey
Zherui Zhang, Rongtao Xu, Jie Zhou +8
The Transformer architecture has achieved significant success in natural language processing, motivating its adaptation to computer vision tasks. Unlike convolutional neural networ…
Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition
Changwei Wang, Shunpeng Chen, Yukun Song +11
Visual Place Recognition (VPR) is aimed at predicting the location of a query image by referencing a database of geotagged images. For VPR task, often fewer discriminative local re…
SkinFormer: Learning Statistical Texture Representation with Transformer for Skin Lesion Segmentation
Rongtao Xu, Changwei Wang, Jiguang Zhang +3
Accurate skin lesion segmentation from dermoscopic images is of great importance for skin cancer diagnosis. However, automatic segmentation of melanoma remains a challenging task b…
HCF-Net: Hierarchical Context Fusion Network for Infrared Small Object Detection
Shibiao Xu, ShuChen Zheng, Wenhao Xu +6
Infrared small object detection is an important computer vision task involving the recognition and localization of tiny objects in infrared images, which usually contain only a few…