7 citations · 29 across the 9 of their papers we have counts for
9 papers
Support-set based Multi-modal Representation Enhancement for Video Captioning
Xiaoya Chen, Jingkuan Song, Pengpeng Zeng +2
Video captioning is a challenging task that necessitates a thorough comprehension of visual scenes. Existing methods follow a typical one-to-one mapping, which concentrates on a li…
Fine-Grained Predicates Learning for Scene Graph Generation
Xinyu Lyu, Lianli Gao, Yuyu Guo +4
The performance of current Scene Graph Generation models is severely hampered by some hard-to-distinguish predicates, e.g., "woman-on/standing on/walking on-beach" or "woman-near/l…
Practical Evaluation of Adversarial Robustness via Adaptive Auto Attack
Ye Liu, Yaya Cheng, Lianli Gao +3
Defense models against adversarial attacks have grown significantly, but the lack of practical evaluation methods has hindered progress. Evaluation can be defined as looking for de…
Unified Multivariate Gaussian Mixture for Efficient Neural Image Compression
Xiaosu Zhu, Jingkuan Song, Lianli Gao +2
Modeling latent variables with priors and hyperpriors is an essential problem in variational image compression. Formally, trade-off between rate and distortion is handled well if p…
One-shot Scene Graph Generation
Yuyu Guo, Jingkuan Song, Lianli Gao +1
As a structured representation of the image content, the visual scene graph (visual relationship) acts as a bridge between computer vision and natural language processing. Existing…
Exploiting long-term temporal dynamics for video captioning
Yuyu Guo, Jingqiu Zhang, Lianli Gao
Automatically describing videos with natural language is a fundamental challenge for computer vision and natural language processing. Recently, progress in this problem has been ac…