14 citations · 14 across the 3 of their papers we have counts for
3 papers
cs.CV2025
PAVE: Patching and Adapting Video Large Language Models
Zhuoming Liu, Yiquan Li, Khoi Duc Nguyen +2
Pre-trained video large language models (Video LLMs) exhibit remarkable reasoning capabilities, yet adapting these models to new tasks involving additional modalities or data types…
cs.CV2021
Learning to Generate Scene Graph from Natural Language Supervision
Yiwu Zhong, Jing Shi, Jianwei Yang +2
Learning from image-text data has demonstrated recent success for many recognition tasks, yet is currently limited to visual features or individual visual concepts such as objects.…
cs.CV2020★ 14 cited
Comprehensive Image Captioning via Scene Graph Decomposition
Yiwu Zhong, Liwei Wang, Jianshu Chen +2
We address the challenging problem of image captioning by revisiting the representation of image scene graph. At the core of our method lies the decomposition of a scene graph into…