554 citations · 853 across the 19 of their papers we have counts for
7 papers · 1 filter
Unsupervised Hashing with Semantic Concept Mining
Rong-Cheng Tu, Xian-Ling Mao, Kevin Qinghong Lin +5
Recently, to improve the unsupervised image retrieval performance, plenty of unsupervised hashing methods have been proposed by designing a semantic similarity matrix, which is bas…
Declaration-based Prompt Tuning for Visual Question Answering
Yuhang Liu, Wei Wei, Daowan Peng +1
In recent years, the pre-training-then-fine-tuning paradigm has yielded immense success on a wide spectrum of cross-modal tasks, such as visual question answering (VQA), in which a…
Context-aware Biaffine Localizing Network for Temporal Sentence Grounding
Daizong Liu, Xiaoye Qu, Jianfeng Dong +5
This paper addresses the problem of temporal sentence grounding (TSG), which aims to identify the temporal boundary of a specific segment from an untrimmed video by a sentence quer…
Spatiotemporal Graph Neural Network based Mask Reconstruction for Video Object Segmentation
Daizong Liu, Shuangjie Xu, Xiao-Yang Liu +3
This paper addresses the task of segmenting class-agnostic objects in semi-supervised setting. Although previous detection based methods achieve relatively good performance, these…
Crowd Counting via Hierarchical Scale Recalibration Network
Zhikang Zou, Yifan Liu, Shuangjie Xu +3
The task of crowd counting is extremely challenging due to complicated difficulties, especially the huge variation in vision scale. Previous works tend to adopt a naive concatenati…
Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation
Wei Wei, Ling Cheng, Xianling Mao +2
Recently, automatic image caption generation has been an important focus of the work on multimodal translation task. Existing approaches can be roughly categorized into two classes…