152 citations · 152 across the 1 of their papers we have counts for
1 paper
Henghui Ding, Chang Liu, Suchen Wang +1
We propose a Vision-Language Transformer (VLT) framework for referring segmentation to facilitate deep interactions among multi-modal information and enhance the holistic understan…