3 citations · 6 across the 3 of their papers we have counts for
3 papers
cs.CV2022★ 3 cited
DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and Grounding
Shilong Liu, Yaoyuan Liang, Feng Li +5
In this paper, we study the problem of visual grounding by considering both phrase extraction and grounding (PEG). In contrast to the previous phrase-known-at-test setting, PEG req…
cs.CV2022★ 2 cited
A Unified Mutual Supervision Framework for Referring Expression Segmentation and Generation
Shijia Huang, Feng Li, Hao Zhang +3
Reference Expression Segmentation (RES) and Reference Expression Generation (REG) are mutually inverse tasks that can be naturally jointly trained. Though recent work has explored…
cs.CV2022★ 1 cited
Multi-View Transformer for 3D Visual Grounding
Shijia Huang, Yilun Chen, Jiaya Jia +1
The 3D visual grounding task aims to ground a natural language description to the targeted object in a 3D scene, which is usually represented in 3D point clouds. Previous works stu…