3 citations · 3 across the 2 of their papers we have counts for
3 papers
cs.CV2023
Progressive Visual Prompt Learning with Contrastive Feature Re-formation
Chen Xu, Yuhan Zhu, Haocheng Shen +4
Prompt learning has been designed as an alternative to fine-tuning for adapting Vision-language (V-L) models to the downstream tasks. Previous works mainly focus on text prompt whi…
cs.CV2022★ 3 cited
Dynamic MDETR: A Dynamic Multimodal Transformer Decoder for Visual Grounding
Fengyuan Shi, Ruopeng Gao, Weilin Huang +1
Multimodal transformer exhibits high capacity and flexibility to align image and text for visual grounding. However, the existing encoder-only grounding framework (e.g., TransVG) s…
cs.CV2021
End-to-End Dense Video Grounding via Parallel Regression
Fengyuan Shi, Weilin Huang, Limin Wang
Video grounding aims to localize the corresponding video moment in an untrimmed video given a language query. Existing methods often address this task in an indirect way, by castin…