84 citations · 97 across the 9 of their papers we have counts for
8 papers
MotionZero:Exploiting Motion Priors for Zero-shot Text-to-Video Generation
Sitong Su, Litao Guo, Lianli Gao +2
Zero-shot Text-to-Video synthesis generates videos based on prompts without any videos. Without motion information from videos, motion priors implied in prompts are vital guidance.…
Continual Referring Expression Comprehension via Dual Modular Memorization
Heng Tao Shen, Cheng Chen, Peng Wang +3
Referring Expression Comprehension (REC) aims to localize an image region of a given object described by a natural-language expression. While promising performance has been demonst…
Visual Commonsense-aware Representation Network for Video Captioning
Pengpeng Zeng, Haonan Zhang, Lianli Gao +3
Generating consecutive descriptions for videos, i.e., Video Captioning, requires taking full advantage of visual representation along with the generation process. Existing video ca…
A Lower Bound of Hash Codes' Performance
Xiaosu Zhu, Jingkuan Song, Yu Lei +2
As a crucial approach for compact representation learning, hashing has achieved great success in effectiveness and efficiency. Numerous heuristic Hamming space metric learning obje…
Support-set based Multi-modal Representation Enhancement for Video Captioning
Xiaoya Chen, Jingkuan Song, Pengpeng Zeng +2
Video captioning is a challenging task that necessitates a thorough comprehension of visual scenes. Existing methods follow a typical one-to-one mapping, which concentrates on a li…
Fine-Grained Predicates Learning for Scene Graph Generation
Xinyu Lyu, Lianli Gao, Yuyu Guo +4
The performance of current Scene Graph Generation models is severely hampered by some hard-to-distinguish predicates, e.g., "woman-on/standing on/walking on-beach" or "woman-near/l…