17 citations · 17 across the 2 of their papers we have counts for
1 paper · 1 filter
Yuan Gong, Andrew Rouditchenko, Alexander H. Liu +4
In this paper, we first extend the recent Masked Auto-Encoder (MAE) model from a single modality to audio-visual multi-modalities. Subsequently, we propose the Contrastive Audio-Vi…