158 citations · 576 across the 42 of their papers we have counts for
1 paper · 2 filters
Alexandros Haliassos, Pingchuan Ma, Rodrigo Mira +2
We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, an…