2 citations · 2 across the 1 of their papers we have counts for
1 paper
Su Zhang, Yi Ding, Ziquan Wei +1
We propose an audio-visual spatial-temporal deep neural network with: (1) a visual block containing a pretrained 2D-CNN followed by a temporal convolutional network (TCN); (2) an a…