6 citations · 6 across the 2 of their papers we have counts for
2 papers
cs.CV2023
Local Compressed Video Stream Learning for Generic Event Boundary Detection
Libo Zhang, Xin Gu, Congcong Li +2
Generic event boundary detection aims to localize the generic, taxonomy-free event boundaries that segment videos into chunks. Existing methods typically require video frames to be…
cs.CV2023★ 6 cited
Text with Knowledge Graph Augmented Transformer for Video Captioning
Xin Gu, Guang Chen, Yufei Wang +3
Video captioning aims to describe the content of videos using natural language. Although significant progress has been made, there is still much room to improve the performance for…