59 citations · 187 across the 17 of their papers we have counts for
6 papers · 2 filters
Confidence-aware Non-repetitive Multimodal Transformers for TextCaps
Zhaokai Wang, Renda Bao, Qi Wu +1
When describing an image, reading text in the visual scene is crucial to understand the key information. Recent work explores the TextCaps task, i.e. image captioning with reading…
Human-centric Spatio-Temporal Video Grounding With Visual Transformers
Zongheng Tang, Yue Liao, Si Liu +5
In this work, we introduce a novel task - Humancentric Spatio-Temporal Video Grounding (HC-STVG). Unlike the existing referring expression tasks in images or videos, by focusing on…
Linguistic Structure Guided Context Modeling for Referring Image Segmentation
Tianrui Hui, Si Liu, Shaofei Huang +4
Referring image segmentation aims to predict the foreground mask of the object referred by a natural language sentence. Multimodal context of the sentence is crucial to distinguish…
Referring Image Segmentation via Cross-Modal Progressive Comprehension
Shaofei Huang, Tianrui Hui, Si Liu +5
Referring image segmentation aims at segmenting the foreground masks of the entities that can well match the description given in the natural language expression. Previous approach…
Recapture as You Want
Chen Gao, Si Liu, Ran He +2
With the increasing prevalence and more powerful camera systems of mobile devices, people can conveniently take photos in their daily life, which naturally brings the demand for mo…
Tree-Structured Policy based Progressive Reinforcement Learning for Temporally Language Grounding in Video
Jie Wu, Guanbin Li, Si Liu +1
Temporally language grounding in untrimmed videos is a newly-raised task in video understanding. Most of the existing methods suffer from inferior efficiency, lacking interpretabil…