2 citations · 2 across the 2 of their papers we have counts for
4 papers
Triple Correlations-Guided Label Supplementation for Unbiased Video Scene Graph Generation
Wenqing Wang, Kaifeng Gao, Yawei Luo +5
Video-based scene graph generation (VidSGG) is an approach that aims to represent video content in a dynamic graph by identifying visual entities and their relationships. Due to th…
Improving Reference-based Distinctive Image Captioning with Contrastive Rewards
Yangjun Mao, Jun Xiao, Dong Zhang +4
Distinctive Image Captioning (DIC) -- generating distinctive captions that describe the unique details of a target image -- has received considerable attention over the last few ye…
TreePrompt: Learning to Compose Tree Prompts for Explainable Visual Grounding
Chenchi Zhang, Jun Xiao, Lei Chen +2
Prompt tuning has achieved great success in transferring the knowledge from large pretrained vision-language models into downstream tasks, and has dominated the performance on visu…
Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language Models
Lin Li, Jun Xiao, Guikun Chen +3
Pretrained vision-language models, such as CLIP, have demonstrated strong generalization capabilities, making them promising tools in the realm of zero-shot visual recognition. Vis…