1 citations · 3 across the 17 of their papers we have counts for
1 paper · 2 filters
Benjamin Robson, Santeri Mentu, Wenshuai Zhao +1
We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, th…