7 citations · 10 across the 5 of their papers we have counts for
4 papers · 1 filter
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
You Qin, Wei Ji, Xinze Lan +5
In the realm of video dialog response generation, the understanding of video content and the temporal nuances of conversation history are paramount. While a segment of current rese…
Towards Complex-query Referring Image Segmentation: A Novel Benchmark
Wei Ji, Li Li, Hao Fei +4
Referring Image Understanding (RIS) has been extensively studied over the past decade, leading to the development of advanced algorithms. However, there has been a lack of research…
In Defense of Clip-based Video Relation Detection
Meng Wei, Long Chen, Wei Ji +2
Video Visual Relation Detection (VidVRD) aims to detect visual relationship triplets in videos using spatial bounding boxes and temporal boundaries. Existing VidVRD methods can be…
Generating Visual Spatial Description via Holistic 3D Scene Understanding
Yu Zhao, Hao Fei, Wei Ji +4
Visual spatial description (VSD) aims to generate texts that describe the spatial relations of the given objects within images. Existing VSD work merely models the 2D geometrical v…