activity
20222025
most citedCollaborative Three-Stream Transformers for Video Captioning

8 citations · 16 across the 8 of their papers we have counts for

collaborators

11 papers

cs.CV2025

Structured Context Learning for Generic Event Boundary Detection

Xin Gu, Congcong Li, Xinyao Wang +5

Generic Event Boundary Detection (GEBD) aims to identify moments in videos that humans perceive as event boundaries. This paper proposes a novel method for addressing this task, ca…

cs.CV20251 cited

High-Fidelity Image Inpainting with Multimodal Guided GAN Inversion

Libo Zhang, Yongsheng Yu, Jiali Yao +1

Generative Adversarial Network (GAN) inversion have demonstrated excellent performance in image inpainting that aims to restore lost or damaged image texture using its unmasked con…

cs.CV2024

Context-Guided Spatio-Temporal Video Grounding

Xin Gu, Heng Fan, Yan Huang +2

Spatio-temporal video grounding (or STVG) task aims at locating a spatio-temporal tube for a specific instance given a text query. Despite advancements, current methods easily suff…

cs.CV2023

Flow-Guided Diffusion for Video Inpainting

Bohai Gu, Yongsheng Yu, Heng Fan +1

Video inpainting has been challenged by complex scenarios like large movements and low-light conditions. Current methods, including emerging diffusion models, face limitations in q…

cs.CV2023

Local Compressed Video Stream Learning for Generic Event Boundary Detection

Libo Zhang, Xin Gu, Congcong Li +2

Generic event boundary detection aims to localize the generic, taxonomy-free event boundaries that segment videos into chunks. Existing methods typically require video frames to be…

cs.CV20238 cited

Collaborative Three-Stream Transformers for Video Captioning

Hao Wang, Libo Zhang, Heng Fan +1

As the most critical components in a sentence, subject, predicate and object require special attention in the video captioning task. To implement this idea, we design a novel frame…