416 citations · 1.1k across the 80 of their papers we have counts for
14 papers · 2 filters
Cross-Modal Collaborative Representation Learning and a Large-Scale RGBT Benchmark for Crowd Counting
Lingbo Liu, Jiaqi Chen, Hefeng Wu +3
Crowd counting is a fundamental yet challenging task, which desires rich information to generate pixel-wise crowd density maps. However, most previous methods only used the limited…
Human-centric Spatio-Temporal Video Grounding With Visual Transformers
Zongheng Tang, Yue Liao, Si Liu +5
In this work, we introduce a novel task - Humancentric Spatio-Temporal Video Grounding (HC-STVG). Unlike the existing referring expression tasks in images or videos, by focusing on…
A Hamiltonian Monte Carlo Method for Probabilistic Adversarial Attack and Learning
Hongjun Wang, Guanbin Li, Xiaobai Liu +1
Although deep convolutional neural networks (CNNs) have demonstrated remarkable performance on multiple computer vision tasks, researches on adversarial learning have shown that de…
Linguistic Structure Guided Context Modeling for Referring Image Segmentation
Tianrui Hui, Si Liu, Shaofei Huang +4
Referring image segmentation aims to predict the foreground mask of the object referred by a natural language sentence. Multimodal context of the sentence is crucial to distinguish…
Referring Image Segmentation via Cross-Modal Progressive Comprehension
Shaofei Huang, Tianrui Hui, Si Liu +5
Referring image segmentation aims at segmenting the foreground masks of the entities that can well match the description given in the natural language expression. Previous approach…
Reinforcement Learning for Weakly Supervised Temporal Grounding of Natural Language in Untrimmed Videos
Jie Wu, Guanbin Li, Xiaoguang Han +1
Temporal grounding of natural language in untrimmed videos is a fundamental yet challenging multimedia task facilitating cross-media visual content retrieval. We focus on the weakl…