4 papers
Human-centric Spatio-Temporal Video Grounding With Visual Transformers
Zongheng Tang, Yue Liao, Si Liu +5
In this work, we introduce a novel task - Humancentric Spatio-Temporal Video Grounding (HC-STVG). Unlike the existing referring expression tasks in images or videos, by focusing on…
Unsupervised Sketch-to-Photo Synthesis
Runtao Liu, Qian Yu, Stella Yu
Humans can envision a realistic photo given a free-hand sketch that is not only spatially imprecise and geometrically distorted but also without colors and visual details. We study…
SketchyScene: Richly-Annotated Scene Sketches
Changqing Zou, Qian Yu, Ruofei Du +6
We contribute the first large-scale dataset of scene sketches, SketchyScene, with the goal of advancing research on sketch understanding at both the object and scene level. The dat…
Large Scale Scene Text Verification with Guided Attention
Dafang He, Yeqing Li, Alexander Gorban +5
Many tasks are related to determining if a particular text string exists in an image. In this work, we propose a new framework that learns this task in an end-to-end way. The frame…