11 citations · 25 across the 4 of their papers we have counts for
4 papers
COSA: Concatenated Sample Pretrained Vision-Language Foundation Model
Sihan Chen, Xingjian He, Handong Li +3
Due to the limited scale and quality of video-text training corpus, most vision-language foundation models employ image-text datasets for pretraining and primarily focus on modelin…
Exploiting Spatial-Temporal Semantic Consistency for Video Scene Parsing
Xingjian He, Weining Wang, Zhiyong Xu +3
Compared with image scene parsing, video scene parsing introduces temporal information, which can effectively improve the consistency and accuracy of prediction. In this paper, we…
Global-Local Propagation Network for RGB-D Semantic Segmentation
Sihan Chen, Xinxin Zhu, Wei Liu +2
Depth information matters in RGB-D semantic segmentation task for providing additional geometric information to color images. Most existing methods exploit a multi-stage fusion str…
Non-Autoregressive Image Captioning with Counterfactuals-Critical Multi-Agent Learning
Longteng Guo, Jing Liu, Xinxin Zhu +3
Most image captioning models are autoregressive, i.e. they generate each word by conditioning on previously generated words, which leads to heavy latency during inference. Recently…