2 citations · 4 across the 4 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2023
Artificial-Spiking Hierarchical Networks for Vision-Language Representation Learning
Yeming Chen, Siyu Zhang, Yaoru Sun +2
With the success of self-supervised learning, multimodal foundation models have rapidly adapted a wide range of downstream tasks driven by vision and language (VL) pretraining. Sta…
cs.CV2023★ 1 cited
LOIS: Looking Out of Instance Semantics for Visual Question Answering
Siyu Zhang, Yeming Chen, Yaoru Sun +3
Visual question answering (VQA) has been intensively studied as a multimodal task that requires effort in bridging vision and language to infer answers correctly. Recent attempts h…
cs.CV2023★ 2 cited
Pixel Difference Convolutional Network for RGB-D Semantic Segmentation
Jun Yang, Lizhi Bai, Yaoru Sun +3
RGB-D semantic segmentation can be advanced with convolutional neural networks due to the availability of Depth data. Although objects cannot be easily discriminated by just the 2D…