152 citations · 168 across the 3 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition
Xiaodan Hu, Chuhang Zou, Suchen Wang +2
Recent video action recognition methods have shown excellent performance by adapting large-scale pre-trained language-image models to the video domain. However, language models con…
cs.CV2022★ 152 cited
VLT: Vision-Language Transformer and Query Generation for Referring Segmentation
Henghui Ding, Chang Liu, Suchen Wang +1
We propose a Vision-Language Transformer (VLT) framework for referring segmentation to facilitate deep interactions among multi-modal information and enhance the holistic understan…
cs.CV2021★ 16 cited
Vision-Language Transformer and Query Generation for Referring Segmentation
Henghui Ding, Chang Liu, Suchen Wang +1
In this work, we address the challenging task of referring segmentation. The query expression in referring segmentation typically indicates the target object by describing its rela…