activity
20222024
most citedText with Knowledge Graph Augmented Transformer for Video Captioning

6 citations · 8 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2024

Context-Guided Spatio-Temporal Video Grounding

Xin Gu, Heng Fan, Yan Huang +2

Spatio-temporal video grounding (or STVG) task aims at locating a spatio-temporal tube for a specific instance given a text query. Despite advancements, current methods easily suff…

cs.CV2023

Local Compressed Video Stream Learning for Generic Event Boundary Detection

Libo Zhang, Xin Gu, Congcong Li +2

Generic event boundary detection aims to localize the generic, taxonomy-free event boundaries that segment videos into chunks. Existing methods typically require video frames to be…

cs.CV20231 cited

Unsupervised Domain Adaptive Detection with Network Stability Analysis

Wenzhang Zhou, Heng Fan, Tiejian Luo +1

Domain adaptive detection aims to improve the generality of a detector, learned from the labeled source domain, on the unlabeled target domain. In this work, drawing inspiration fr…

cs.CV20236 cited

Text with Knowledge Graph Augmented Transformer for Video Captioning

Xin Gu, Guang Chen, Yufei Wang +3

Video captioning aims to describe the content of videos using natural language. Although significant progress has been made, there is still much room to improve the performance for…

cs.SE20221 cited

AUGER: Automatically Generating Review Comments with Pre-training Models

Lingwei Li, Li Yang, Huaxi Jiang +5

Code review is one of the best practices as a powerful safeguard for software quality. In practice, senior or highly skilled reviewers inspect source code and provide constructive…

cs.CV2022

High-Fidelity Image Inpainting with GAN Inversion

Yongsheng Yu, Libo Zhang, Heng Fan +1

Image inpainting seeks a semantically consistent way to recover the corrupted image in the light of its unmasked content. Previous approaches usually reuse the well-trained GAN as…