1 paper
Shucong Zhang, Titouan Parcollet, Rogier van Dalen +1
Self-supervised learning (SSL) models usually require weeks of pre-training with dozens of high-end GPUs. These models typically have a multi-headed self-attention (MHSA) context e…