6 citations · 8 across the 7 of their papers we have counts for
7 papers
Context-Guided Spatio-Temporal Video Grounding
Xin Gu, Heng Fan, Yan Huang +2
Spatio-temporal video grounding (or STVG) task aims at locating a spatio-temporal tube for a specific instance given a text query. Despite advancements, current methods easily suff…
Local Compressed Video Stream Learning for Generic Event Boundary Detection
Libo Zhang, Xin Gu, Congcong Li +2
Generic event boundary detection aims to localize the generic, taxonomy-free event boundaries that segment videos into chunks. Existing methods typically require video frames to be…
Unsupervised Domain Adaptive Detection with Network Stability Analysis
Wenzhang Zhou, Heng Fan, Tiejian Luo +1
Domain adaptive detection aims to improve the generality of a detector, learned from the labeled source domain, on the unlabeled target domain. In this work, drawing inspiration fr…
Text with Knowledge Graph Augmented Transformer for Video Captioning
Xin Gu, Guang Chen, Yufei Wang +3
Video captioning aims to describe the content of videos using natural language. Although significant progress has been made, there is still much room to improve the performance for…
AUGER: Automatically Generating Review Comments with Pre-training Models
Lingwei Li, Li Yang, Huaxi Jiang +5
Code review is one of the best practices as a powerful safeguard for software quality. In practice, senior or highly skilled reviewers inspect source code and provide constructive…
High-Fidelity Image Inpainting with GAN Inversion
Yongsheng Yu, Libo Zhang, Heng Fan +1
Image inpainting seeks a semantically consistent way to recover the corrupted image in the light of its unmasked content. Previous approaches usually reuse the well-trained GAN as…