1 paper · 1 filter
Juan Yeo, Jinkwan Jang, Kyubyung Chae +2
Recent studies show that pretrained vision models can boost performance in audio downstream tasks. To enhance the performance further, an additional pretraining stage with large sc…