3 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.CV2024
Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark
Zeyu Xi, Ge Shi, Xuefen Li +5
Despite the recent emergence of video captioning models, how to generate the text description with specific entity names and fine-grained actions is far from being solved, which ho…
cs.CV2022★ 3 cited
Learning to Compose Diversified Prompts for Image Emotion Classification
Sinuo Deng, Lifang Wu, Ge Shi +3
Contrastive Language-Image Pre-training (CLIP) represents the latest incarnation of pre-trained vision-language models. Although CLIP has recently shown its superior power on a wid…