1 citations · 1 across the 2 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2023
Artificial-Spiking Hierarchical Networks for Vision-Language Representation Learning
Yeming Chen, Siyu Zhang, Yaoru Sun +2
With the success of self-supervised learning, multimodal foundation models have rapidly adapted a wide range of downstream tasks driven by vision and language (VL) pretraining. Sta…
cs.CV2023★ 1 cited
LOIS: Looking Out of Instance Semantics for Visual Question Answering
Siyu Zhang, Yeming Chen, Yaoru Sun +3
Visual question answering (VQA) has been intensively studied as a multimodal task that requires effort in bridging vision and language to infer answers correctly. Recent attempts h…