83 citations · 87 across the 8 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2023★ 2 cited
COSA: Concatenated Sample Pretrained Vision-Language Foundation Model
Sihan Chen, Xingjian He, Handong Li +3
Due to the limited scale and quality of video-text training corpus, most vision-language foundation models employ image-text datasets for pretraining and primarily focus on modelin…
cs.CV2018
Dual Attention Network for Scene Segmentation
Jun Fu, Jing Liu, Haijie Tian +4
In this paper, we address the scene segmentation task by capturing rich contextual dependencies based on the selfattention mechanism. Unlike previous works that capture contexts by…
cs.CV2017★ 83 cited
Stacked Deconvolutional Network for Semantic Segmentation
Jun Fu, Jing Liu, Yuhang Wang +1
Recent progress in semantic segmentation has been driven by improving the spatial resolution under Fully Convolutional Networks (FCNs). To address this problem, we propose a Stacke…