activity
20172022
most citedFrom Deterministic to Generative: Multi-Modal Stochastic RNNs for Video Captioning

16 citations · 34 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV20221 cited

Fine-Grained Predicates Learning for Scene Graph Generation

Xinyu Lyu, Lianli Gao, Yuyu Guo +4

The performance of current Scene Graph Generation models is severely hampered by some hard-to-distinguish predicates, e.g., "woman-on/standing on/walking on-beach" or "woman-near/l…

cs.CV20225 cited

One-shot Scene Graph Generation

Yuyu Guo, Jingkuan Song, Lianli Gao +1

As a structured representation of the image content, the visual scene graph (visual relationship) acts as a bridge between computer vision and natural language processing. Existing…

cs.CV202211 cited

Exploiting long-term temporal dynamics for video captioning

Yuyu Guo, Jingqiu Zhang, Lianli Gao

Automatically describing videos with natural language is a fundamental challenge for computer vision and natural language processing. Recently, progress in this problem has been ac…

cs.CV20221 cited

Relation Regularized Scene Graph Generation

Yuyu Guo, Lianli Gao, Jingkuan Song +4

Scene graph generation (SGG) is built on top of detected objects to predict object pairwise visual relations for describing the image content abstraction. Existing works have revea…

cs.CV2021

From General to Specific: Informative Scene Graph Generation via Balance Adjustment

Yuyu Guo, Lianli Gao, Xuanhan Wang +5

The scene graph generation (SGG) task aims to detect visual relationship triplets, i.e., subject, predicate, object, in an image, providing a structural vision layout for scene und…

cs.CV201716 cited

From Deterministic to Generative: Multi-Modal Stochastic RNNs for Video Captioning

Jingkuan Song, Yuyu Guo, Lianli Gao +3

Video captioning in essential is a complex natural process, which is affected by various uncertainties stemming from video content, subjective judgment, etc. In this paper we build…