95 citations · 143 across the 11 of their papers we have counts for
5 papers · 1 filter
Towards Complex-query Referring Image Segmentation: A Novel Benchmark
Wei Ji, Li Li, Hao Fei +4
Referring Image Understanding (RIS) has been extensively studied over the past decade, leading to the development of advanced algorithms. However, there has been a lack of research…
NExT-GPT: Any-to-Any Multimodal LLM
Shengqiong Wu, Hao Fei, Leigang Qu +2
While recently Multimodal Large Language Models (MM-LLMs) have made exciting strides, they mostly fall prey to the limitation of only input-side multimodal understanding, without t…
In Defense of Clip-based Video Relation Detection
Meng Wei, Long Chen, Wei Ji +2
Video Visual Relation Detection (VidVRD) aims to detect visual relationship triplets in videos using spatial bounding boxes and temporal boundaries. Existing VidVRD methods can be…
Panoptic Scene Graph Generation with Semantics-Prototype Learning
Li Li, Wei Ji, Yiming Wu +4
Panoptic Scene Graph Generation (PSG) parses objects and predicts their relationships (predicate) to connect human language and visual scenes. However, different language preferenc…
Visually-Prompted Language Model for Fine-Grained Scene Graph Generation in an Open World
Qifan Yu, Juncheng Li, Yu Wu +3
Scene Graph Generation (SGG) aims to extract <subject, predicate, object> relationships in images for vision understanding. Although recent works have made steady progress on SGG,…