24 citations · 30 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 6 cited
Text with Knowledge Graph Augmented Transformer for Video Captioning
Xin Gu, Guang Chen, Yufei Wang +3
Video captioning aims to describe the content of videos using natural language. Although significant progress has been made, there is still much room to improve the performance for…
cs.CV2023★ 24 cited
DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training
Wei Li, Linchao Zhu, Longyin Wen +1
Large-scale pre-trained multi-modal models (e.g., CLIP) demonstrate strong zero-shot transfer capability in many discriminative tasks. Their adaptation to zero-shot image-condition…